Tools
Is AI Pronunciation Feedback Accurate for Korean?
Accuracy depends on what the scorer compares, and most of them compare the wrong thing. Korean changes its sounds when letters meet, so a correctly pronounced word looks nothing like its spelling. A scorer that checks spelling will fail you for getting it right.
We know because we built one and it did exactly that to our own student.
What went wrong, specifically
Speech recognition writes down what it heard. The better a learner pronounces Korean, the further that transcript moves from the written form of the word. Our first scorer compared the transcript to the spelling.
많다
man-ta
many (pronounced man-ta)
있다
it-tta
to be (pronounced it-tta)
식당
sik-ttang
restaurant (pronounced sik-ttang)
Three words, all said correctly, all failed. The scorer was not being harsh. It was measuring the wrong thing, and what it punished was pronunciation that had got better.
The fix was to compare sound to sound rather than sound to spelling. On that set, passing attempts went from 200 to 175: 28 dropped out and all of them were genuine errors, while three that had been failing started passing. Figures are from the 2026-09-13 run.
Why being generous is not the fix
The obvious repair is to loosen the threshold, and it is the wrong one. A scorer that accepts anything close will pass 병원 when the student said 평원, and 책 when they said 체.
Both halves are needed: it has to compare sound, and it has to be strict about sound. Loose-and-spelling-based is the worst combination, and it is the easiest one to build.
The question to ask any tool
- Does it fail you on words you know you said right? If yes, it is comparing spelling. This shows up on words with a final consonant meeting a new one.
- Does it pass you on words you know you fumbled? If yes, the threshold is doing the work instead of the comparison.
- Does it tell you what was wrong, or only give a number? A number you cannot act on trains anxiety rather than pronunciation.
The thing no scorer does
There is a category it cannot reach, and it is not a technical limit. A scorer marks the word you were asked to say. It has no opinion about the word you should have said instead.
It also cannot hear that a recording failed rather than a student. We had to build that distinction in separately, because a clipped recording stored as a wrong answer makes a teacher misread the student, and the student starts doubting a pronunciation that was fine.
So can it replace a person?
For repetition, it is better than a person, because it never gets tired and it is available at two in the morning. That is a real advantage and it is worth using.
What it does not do is change what happens next. A scorer marks the attempt. A teacher hears the attempt and picks the next thing because of it, which is a different job and the reason lessons are not just a list of exercises.
The useful arrangement is not one or the other. It is a tool for the reps and a person for the direction, and neither one doing the other's job.
What to do this week
- Take three words with a final consonant running into the next syllable and check what your tool gives you when you say them correctly.
- If it fails them, stop treating its score as information about your mouth.
- Use it for volume, not for judgment. The reps are real even when the number is not.
- Get the direction from somebody who can hear the attempt and choose what comes next.
The short version
Accurate is the wrong question. Ask what it compares. If it compares spelling it will punish the pronunciation you worked hardest for, and we know that because we shipped that bug into our own app before we caught it.
The part a score cannot do
A number tells you the attempt was off. It does not decide what you practise next. That decision is what a 1:1 session is for. How the hour is structured is explained there.
See the first session →