Report · July 24, 2026

Transcription accuracy, measured — not marketed

Every transcription service claims “industry-leading accuracy.” Almost none publish audio you can check. We ran 22 minutes of public audio with verifiable reference transcripts — a real four-person meeting, clean read speech, and Nigerian-accented English — through our live product and two leading engines, and we kept everything: the file manifests, the references, the raw outputs, and the scoring script. Want to check our numbers? Email us for the reproduction pack.

6.43%

ConvertAudioToText avg WER — best overall

2 / 2

African-accented English sets won outright

100%

public audio with verifiable reference transcripts

Results — word error rate (lower is better)

AudioSourceCATTDeepgram nova-3Whisper v3-turbo
Nigerian English — set 14.5 min · 537 reference wordsOpenSLR 70 (CC BY-SA 4.0)5.15%5.51%6.25%
Nigerian English — set 24 min · 491 reference wordsOpenSLR 70 (CC BY-SA 4.0)5.63%6.24%9.46%
Real meeting (4 speakers)10 min · 1105 reference wordsAMI Corpus ES2004a (CC BY 4.0)23.59%23.25%24.02%
Clean speech — reader 10892.6 min · 410 reference wordsLibriSpeech test-clean0.98%1.22%1.22%
Clean speech — reader 1211.5 min · 268 reference wordsLibriSpeech test-clean0.74%0.00%3.31%
Clean speech — reader 45072.1 min · 240 reference wordsLibriSpeech test-clean2.50%2.50%2.50%
Average (all six files)6.43%6.45%7.79%

Three honest findings

  1. Accents are where engines separate. On Nigerian-accented English (OpenSLR 70), our pipeline won both sets outright, and Whisper — the open model many budget tools are built on — degraded the most (up to 9.46% WER). If your speakers aren’t American newsreaders, accent performance matters more than any headline number.
  2. On real meetings, top engines are nearly tied. All three landed within one point on a genuine four-person meeting (AMI ES2004a) — which includes overlapping speech and verbal fillers that inflate everyone’s WER. At this level, accuracy is table stakes: what separates products is what they do with the transcript — speaker labels, summaries, action items, and exports.
  3. Benchmarking is easy to get wrong — we did, briefly. Our first scoring pass counted our own export formatting (headers, timestamps, speaker labels) as transcription errors and made our product look 3–4 points worse than it is. We found it by diffing outputs word-by-word and fixed the scorer — and the corrected scorer ships in the reproduction pack, so you can check our math.

Method

Check our numbers

The complete pack — audio manifests, reference transcripts, every raw engine output, and the scoring script — is available to anyone who asks. Run it, and if you get different numbers, tell us and we’ll publish the correction.

Request the reproduction pack →

Try the engine that won on your own audio

3 free transcriptions a day, up to 10 minutes each — speaker labels, AI summary and action items included. No credit card.

Transcribe a file free