Report · July 24, 2026
Transcription accuracy, measured — not marketed
Every transcription service claims “industry-leading accuracy.” Almost none publish audio you can check. We ran 22 minutes of public audio with verifiable reference transcripts — a real four-person meeting, clean read speech, and Nigerian-accented English — through our live product and two leading engines, and we kept everything: the file manifests, the references, the raw outputs, and the scoring script. Want to check our numbers? Email us for the reproduction pack.
6.43%
ConvertAudioToText avg WER — best overall
2 / 2
African-accented English sets won outright
100%
public audio with verifiable reference transcripts
Results — word error rate (lower is better)
| Audio | Source | CATT | Deepgram nova-3 | Whisper v3-turbo |
|---|---|---|---|---|
| Nigerian English — set 14.5 min · 537 reference words | OpenSLR 70 (CC BY-SA 4.0) | 5.15% | 5.51% | 6.25% |
| Nigerian English — set 24 min · 491 reference words | OpenSLR 70 (CC BY-SA 4.0) | 5.63% | 6.24% | 9.46% |
| Real meeting (4 speakers)10 min · 1105 reference words | AMI Corpus ES2004a (CC BY 4.0) | 23.59% | 23.25% | 24.02% |
| Clean speech — reader 10892.6 min · 410 reference words | LibriSpeech test-clean | 0.98% | 1.22% | 1.22% |
| Clean speech — reader 1211.5 min · 268 reference words | LibriSpeech test-clean | 0.74% | 0.00% | 3.31% |
| Clean speech — reader 45072.1 min · 240 reference words | LibriSpeech test-clean | 2.50% | 2.50% | 2.50% |
| Average (all six files) | 6.43% | 6.45% | 7.79% | |
Three honest findings
- Accents are where engines separate. On Nigerian-accented English (OpenSLR 70), our pipeline won both sets outright, and Whisper — the open model many budget tools are built on — degraded the most (up to 9.46% WER). If your speakers aren’t American newsreaders, accent performance matters more than any headline number.
- On real meetings, top engines are nearly tied. All three landed within one point on a genuine four-person meeting (AMI ES2004a) — which includes overlapping speech and verbal fillers that inflate everyone’s WER. At this level, accuracy is table stakes: what separates products is what they do with the transcript — speaker labels, summaries, action items, and exports.
- Benchmarking is easy to get wrong — we did, briefly. Our first scoring pass counted our own export formatting (headers, timestamps, speaker labels) as transcription errors and made our product look 3–4 points worse than it is. We found it by diffing outputs word-by-word and fixed the scorer — and the corrected scorer ships in the reproduction pack, so you can check our math.
Method
- Audio (22.7 min, 3,051 reference words): AMI Meeting Corpus ES2004a first 10 minutes (CC BY 4.0, gold multi-speaker transcript); LibriSpeech test-clean, three readers (public-domain audiobooks with exact texts); OpenSLR 70 crowdsourced Nigerian English, 80 utterances (CC BY-SA 4.0).
- Engines: ConvertAudioToText (live product) — files uploaded through the same product customers use, no special configuration; Deepgram nova-3 (direct API); Whisper large-v3-turbo (direct API — the open model family TurboScribe builds on).
- Scoring: word error rate via
jiwer, after normalization applied identically to all engines: lowercase, punctuation stripped, common contractions expanded, export formatting removed. The normalization code is included in the reproduction pack. - Limits, stated plainly: six files is a spot check, not a census; AMI reference transcripts include disfluencies that raise absolute WER for every engine; we tested English only; and we are, obviously, a participant — which is exactly why we share the full pack with anyone who wants to re-run it adversarially.
Check our numbers
The complete pack — audio manifests, reference transcripts, every raw engine output, and the scoring script — is available to anyone who asks. Run it, and if you get different numbers, tell us and we’ll publish the correction.
Request the reproduction pack →Try the engine that won on your own audio
3 free transcriptions a day, up to 10 minutes each — speaker labels, AI summary and action items included. No credit card.
Transcribe a file free