Eventually, most people trying to learn Thai hit a wall. That’s a normal feeling for such a difficult language, but it’s frustrating nonetheless.
At a certain point, learners usually pick one of two options:
But as AI has gotten increasingly powerful, a third option has started to present itself. There’s now a reasonable body of research that suggests pairing a human tutor with AI-driven practice beats either option used alone, with the gap widest for the learners who struggle most.
Evidence for this claim comes from Carnegie Mellon University, where researchers have run a series of studies measuring the efficacy of what they call hybrid human-AI tutoring.
The setup is specific. Students work through problems on an adaptive practice platform, with the software choosing what comes next based on how they are answering. A tutor joins the session remotely over video and works with about four students at a time, moving between them as they go. When a student is stuck, the tutor explains. When a student is doing well, the tutor says so. In the most recent study the tutors were graduate and undergraduate students, and sessions ran twice a week for around thirty minutes of practice.
Their largest study, covering 635 middle-school students, showed:
Source: Carnegie Mellon University, 2026
Those numbers come from their most recent work, but they are not a one-off result. The same group had already published three earlier studies, covering 585 students across three different schools, and found the same direction every time: students in the hybrid model gained more than comparable students, and the ones who gained the most were the ones who had been performing worst.
The pattern underneath those numbers is the one worth paying attention to. AI practice on its own was least effective for the students who needed it most. Lower-performing students left alone with the software disengaged, and what brought them back was not a better algorithm. It was a person noticing.
The researchers describe the human contribution as instructional and motivational support. In plainer terms: the explanation, and the encouragement. Not the drilling.
This is the case against learning a language from an app alone, and it is not a marketing argument. It is what happened when researchers measured it.
A human tutor and an AI learning tool have the same goal: help you learn. They are good at almost opposite parts of it.
AI is better at volume and availability. It will not get tired drilling the same word forty times over, and it is there at whatever hour you have twenty minutes free. It also never makes you feel stupid, which matters more than it sounds. The Brookings Institution, a public policy research group in Washington, notes that AI tutors respond with “infinite patience and non-judgmental support”, and learners will admit confusion to software that they would quietly hide from a person.
Humans are better at judgment. Explaining why something is wrong rather than only that it was. Deciding which of your errors is worth fixing this week and which can wait. And maybe most important of all, being someone you have to show up for. When motivation dips, a person on the other end is what keeps the practice happening at all.
Brookings calls the combination of the two “human-AI hybrid vigor”, while cautioning that many claims about AI in education have run ahead of the evidence supporting them.
Thai makes both halves of this equation harder.
For one, it is very hard for learners to hear their own mistakes in a tonal language. Someone saying a high tone as a falling tone will not usually catch themselves making an error, because they hear themselves saying the word they intended. Catching these mistakes takes a native ear, which means if a learner is practicing on their own, they can practice the wrong thing for months without ever knowing. This is a diagnosis problem, and it is the clearest reason a human belongs in the process at all.
Second, AI software stands on weaker ground with Thai. Speech technology is built and measured largely around well-resourced languages, and Thai is not one of them. The main public benchmark for speech recognition evaluates English as its primary track, along with German, French, Italian, Spanish and Portuguese. Thai is not on the list.
There’s also far less Thai text and speech available for AI models to train on compared to widely spoken languages like English or Mandarin. This doesn’t mean AI is useless for learning Thai, but it does mean your typical AI software is less reliable than it is for learning other languages.
All of these problems compound for Thai learners. If you cannot reliably catch your own errors, and software is less reliable at catching them for you, a human needs to enter the conversation.
But it cuts both ways. Diagnosis alone changes very little, and simply being told that your high tones keep falling will not fix your ability to speak them. Fixing that kind of error takes hundreds of repetitions, and paying a human by the hour to sit through all those reps gets expensive quickly.
This is why pairing a human tutor with AI practice works so well for Thai. Neither half is perfect on its own, but together they give a Thai learner exactly what they need: someone who can hear what is going wrong, and a tool that will drill it as many times as it takes.
We build speakthai.ai around these core ideas. Everything starts with a 15-minute assessment where a human tutor establishes your baseline. From there, four 30-minute tutoring sessions per month help you diagnose what’s wrong and decide where to focus next. In between those lessons, our app builds customized practice based on your tutor’s feedback, so all of the repetition you do in the app is aimed at your specific learning profile, not a generic syllabus.
You can get a long way, and plenty of people do. What the research suggests is that you are most likely to stall early, and that pronunciation errors are the ones least likely to correct themselves. If you go app-only, find some way to get a native speaker to listen to you from time to time.
On current evidence, not for everything. AI does volume, patience and availability better. Humans do diagnosis, judgment and accountability better. The studies that combined them beat either one alone.
It is an argument that arrives at the same place as our product, which is worth saying plainly rather than pretending otherwise. The research is public, it is linked below, it was not commissioned by us and it is not about us. Read it and decide for yourself.
Four 30-minute lessons a month with a Thai tutor, and an app that turns what they flag into your daily practice.
The Carnegie Mellon studies cited here were quasi-experimental rather than randomized trials, run with middle-school students in mathematics, using adaptive maths software rather than the conversational AI most language learners use today.
We cite them because the mechanism they identify is general, not because their effect sizes transfer to adults learning Thai. We are not aware of an equivalent study in adult language learning. If one exists, we would like to read it: support@speakthai.ai.