Speech to Text Software

Speech to text software, without the guesswork

Upload an audio or video file, or paste a link to a recording. You get a speaker-labeled transcript with timestamps in 99 languages, plus summaries, action items and clean exports. The first 10 minutes of any file are free with no account.

Prefer to dictate rather than upload a file? Try live voice typing in your browser

First 10 minutes of any file free without an account. Free accounts: 10 transcription minutes, one time. Pro: unlimited, from $9.99/mo billed annually.

How to convert speech to text

  1. 1
    Bring your audioUpload an MP3, WAV, M4A, MP4, MOV or WebM file, or paste a public link to a recording. Audio is extracted from video automatically, so there is nothing to convert first.
  2. 2
    Name the spoken languageChoose the language before you start. Naming it up front lets the model optimise for the right sound set instead of guessing, which is where most wrong-language transcripts come from. 99 languages are supported, with output in the native script.
  3. 3
    Read, edit and exportYour transcript arrives with speaker labels and timestamps. Reading and copying it is free. Paid plans add downloads in TXT, PDF, DOCX, SRT, VTT and JSON, plus AI summaries, action items and decisions.

What you get back

Speaker labels that separate who spoke when, across the whole recording
Word-level timestamps you can jump to, so every quote stays checkable
99 languages with output in the native script, not romanised text
11 task-tuned AI templates for interviews, podcasts, lectures, meetings and research
Action items, decisions and a ready-to-send email draft pulled from the transcript, written in the transcript's own language
A real editor: fix a word, leave a comment, keep the version history

How accurate it is, measured

We publish our own numbers instead of a marketing percentage. Across a six-file public test set — LibriSpeech clean speech, Nigerian-accented English from OpenSLR 70, and a four-speaker AMI meeting — the live product averaged 6.43% word error rate, measured on July 24, 2026. Meeting audio with crosstalk is far harder than clean speech, and the report breaks out every file, including the ones where we do worst.

Read the full benchmark

What people say about the transcripts

How easy it makes working with other tools, it's speed and it's efficiency. The interface is well coded and visually appealing. This is an incredible tool.
David E. · Marketing Manager
Best transcription tool I've ever used. The accuracy is amazing and the turnaround time is incredibly fast. Highly recommend!
Emily K. · Journalist
The multilingual support is what sold me. Transcribes French and Spanish recordings just as well as English.
Sofia G. · Translator

The same engine, over an API

If you are wiring speech to text into your own product, the Developer and Business plans expose a REST API: create a key in the dashboard, send a file or a URL, then poll for the result and read back the transcript with timestamps and speaker labels. Webhook callbacks are on the Business plan.

See API plans

Frequently asked questions

Is speech to text free?
The first 10 minutes of any file are transcribed free with no account. A free account adds 10 transcription minutes, given once rather than refilled, spendable at up to 3 files a day of 10 minutes each. Reading and copying the text stays free; file downloads and unlimited length come with a paid plan from $9.99/mo billed annually.
How accurate is speech to text?
It depends on the audio far more than on the tool. On clean, close-mic speech our measured word error rate sits near 1%. On a four-speaker meeting with crosstalk it climbs past 20%. We publish the whole test set on our accuracy benchmark page instead of quoting one number. Expect to proofread names, jargon and anything spoken over background noise.
Which languages are supported?
99 languages, with the transcript written in the language's own script — Arabic in Arabic, Hindi in Devanagari, Japanese in kana and kanji. Pick the language before you upload. AI summaries and action items are generated in the same language as the transcript, not translated into English.
Can I convert speech from a video?
Yes. Upload MP4, MOV, WebM or MKV and the audio track is extracted for you — there is no need to convert the video first. You can also paste a public link to a recording instead of uploading a file.
Does it identify who is speaking?
Yes. Speaker diarization runs on every transcript: it separates the voices and labels them Speaker 1, Speaker 2 and so on. It does not know their real names — you rename a speaker once in the editor and the label updates across the whole transcript.
What happens to my audio afterwards?
Free-plan files are deleted within 7 days. Paid plans let you choose the retention window — 30 days by default, up to keeping the file indefinitely — and you can delete any file yourself at any time.

Turn your first recording into text

Nothing to install, no card. The first 10 minutes are on us.

Start transcribing