How a score is produced
Every answer goes through the same pipeline, the same way each time:
- You answer. You speak or write a response to an exam-style task.
- Your speech is measured. For speaking, your audio is transcribed and measured for acoustic features — speaking rate, pauses, rhythm and clarity.
- It's graded on the official rubric. The answer is assessed criterion by criterion using the real exam's rubric: task response, coherence, grammar, vocabulary, and pronunciation or fluency.
- It's mapped to the real scale. The result is placed on the actual exam scale — IELTS bands, TOEFL's 1–6 CEFR bands, the Duolingo English Test's 10–160 — with a calibration step so the numbers line up with official scoring.
Then you get the score, a breakdown for each criterion, and sentence-by-sentence feedback.
How we measure it
We don't grade by feel. The scorer is benchmarked against official, examiner-scored sample answers published by the test makers — IELTS partners, ETS (TOEFL), and Duolingo. Each one is a real response with the official band or score attached.
For every sample, we run it through the exact production scoring pipeline and compare Aflo's score to the official one, using the same measures examiners use to check agreement:
- MAE — how far Aflo's score is from the official score, on average.
- Exact and adjacent agreement — how often Aflo lands on the official band, or within one of it.
- QWK (quadratic weighted kappa) — the standard statistic for agreement between two raters.
Every change to the scorer is re-measured against these references and kept only if agreement improves. The benchmark runs the same code that runs in production, so the numbers reflect what you actually get.
How close it lands
Measured against the official examiner scores:
| What | Agreement with official scores |
|---|---|
| IELTS (writing & speaking) | within about 0.5–0.7 of a band, on average |
| Duolingo English Test (speaking) | QWK ≈ 0.90 — strong agreement |
| Fluency & delivery | ρ ≈ 0.8 vs human ratings |
| Pronunciation (per sound) | ≈ 0.7 vs human pronunciation ratings |
On IELTS, that is close to how closely two trained examiners agree with each other.
Fluency and pronunciation, grounded in real speech
Speaking isn't graded from the transcript alone — Aflo listens to two different things in your audio.
Fluency & delivery — the flow of your speech: pace, pauses, rhythm and clarity. These measurements line up with human fluency ratings at about ρ = 0.8, validated on public research corpora of human-rated learner speech — the ICNALE corpus (Kobe University) and SpeechOcean.
Pronunciation — scored sound by sound. Aflo aligns what you said to the expected pronunciation and measures how closely each phoneme matches, so it doesn't just hand you a number — it tells you which sounds to work on (say, a weak th or r). It agrees with human pronunciation ratings at about ρ ≈ 0.7, and a built-in floor keeps a clear but accented voice from being unfairly marked down.
Ranking well is not the same as scoring correctly
A research corpus tells you whether a measurement ranks speakers correctly. It does not tell you whether the number on the certificate is right. Those are different questions, and we got the second one wrong for a while.
Our fluency measurements ranked speakers well — that is the ρ = 0.8 above. But when we tested them against real IELTS candidates, on a set of 421 answers from 169 recorded mock tests with the awarded band attached (bands 3.5 to 9), the absolute score was badly off: it ran about 1.4 bands too harsh in the 4–6 range, while remaining accurate at 8–9. A genuinely mid-band answer was being told it was a low-band one.
So we re-fitted the curve that turns those measurements into a band, using the real, band-labelled candidate speech rather than the research corpus. We fitted it on one subset and checked it on a held-out 257 answers the fit had never seen:
| On held-out real IELTS speech | Before | After |
|---|---|---|
| Average error per answer | 1.38 bands | 0.99 bands |
| What a band-4 answer scored (median) | 2.78 | 4.01 |
| What a band-6 answer scored (median) | 5.03 | 6.25 |
| What a band-9 answer scored (median) | 8.99 | 8.99 — top of the scale intact |
The new curve is monotone, which means it never changes the order of two speakers — so the ranking accuracy validated on the research corpora still holds. It only fixes where the numbers land.
A score you can act on
A number on its own doesn't help you improve. For every answer, Aflo shows the score for each criterion, then goes sentence by sentence — what's working, what's holding the score down, and how to fix it. You can see exactly why a score is what it is, and what to change to raise it.
Frequently asked questions
What scale does Aflo use?
The real scale for each exam: IELTS bands (0–9), TOEFL's 1–6 CEFR bands, and the Duolingo English Test's 10–160.
What does the score compare against?
Official sample answers that the test makers published with the examiner's own band or score. Aflo runs each one through the production pipeline and compares its score to the official one.
What is QWK?
Quadratic weighted kappa — the standard statistic for how closely two raters agree, used widely in exam-scoring research. Higher is better; around 0.9 means strong agreement.