Free Descript Alternative: Edit Audio by Editing Text, No Signup
The edit-by-text part of Descript, free in your browser. Not the rest of the suite, and we list what is missing.
CATT Editor is a free, browser-only stand-in for one part of Descript: cutting a recording by deleting words, plus one-click removal of the English filler sounds. It needs no account, no install and no upload, and it costs nothing to export. It is not a full Descript replacement, because there is no video suite, no AI voice, no screen recording and no collaboration.
The editor is a separate free app on edit.convertaudiototext.com. It opens in a new tab, asks for nothing, and reads your file straight off your disk.
Works best in a Chromium desktop browser. The first run downloads the speech model, about 200 MB for Base or about 600 MB for Small, then caches it. Everything is held in memory, so short clips work best.
Comparing the full transcription service instead? Read ConvertAudioToText vs Descript
Where the free editor stands in for Descript, and where it does not
You get: cutting by deleting words
Drop in a recording, get text with a timestamp on every word, delete the parts you do not want, and the media is cut to match. Playback skips the cuts straight away.
You get: English filler-sound removal
The editor matches a fixed list of English filler sounds: um, uh, ah, eh, hm, mm, uh-huh and their common spellings. When the transcript holds any of them, a Remove fillers button appears with the count and one click cuts them all. It is a word list rather than a judgement call, so like, so and right stay in, and the list is matched whatever language you spoke: on a non-English recording it can flag a word that carries meaning. Undo puts everything back.
You get: nothing is uploaded
Descript processes your media in its cloud. Here the file is read off your disk, decoded and turned into text on your own hardware, and exported back to your disk. Anonymous product analytics are the only thing we collect, and they carry no file data.
You get: no account and no install
Descript asks you to create an account before the first edit and keeps its full feature set in a desktop app. This is one browser tab: no sign-up, no download, no watermark on the export.
You do not get: the rest of the suite
Descript is a full production tool. It has a multi-track video timeline, AI voice and overdub, screen and studio recording, one-click denoise, and shared projects with comments. None of that exists here, and no free browser tab is going to replace it.
You do not get: publish-grade text on long files
The in-browser model is small, so it struggles with accents, crosstalk, background noise and non-English audio, and your browser memory caps how long a file you can hold. Run long or difficult recordings through ConvertAudioToText, then bring the result into the editor for the cutting.
How it works
Check the trade first
If you need a multi-track video timeline, AI voice, screen recording or a shared project with comments, Descript is the right tool and this is not. If what you want is to cut a recording by deleting words, read on.
Open the editor and drop in a file
MP3, WAV, M4A, MP4, WebM, MOV or MKV. Nothing is uploaded. The first run downloads the speech model, about 200 MB for Base or about 600 MB for Small, then caches it.
Delete what you do not want
The text carries a timestamp on every word, so a deletion cuts the media on the word boundary. The English filler sounds come out in one click.
Export and keep the file
MP4 for video projects, M4A for audio-only ones, written in the browser by ffmpeg.wasm. No watermark, no export credits, no plan to pick.
CATT Editor vs Descript, honestly
Five of these rows go to Descript. We are not going to pretend otherwise.
| Capability | CATT | Descript |
|---|---|---|
| Cost to edit and export | Free, no tier | Metered free tier, paid from about $16/mo |
| Account required | No | Yes, before the first edit |
| Install | None, one browser tab | Browser version is limited, desktop app for the full set |
| Where your media is processed | On your own machine | Uploaded to their cloud |
| Cut by deleting words | Yes | Yes |
| Filler-word removal | Yes, English filler sounds | Yes |
| Multi-track video timeline | No | Yes |
| AI voice and overdub | No | Yes |
| Screen and studio recording | No | Yes |
| Team collaboration and comments | No | Yes |
| Long or difficult recordings | Capped by your browser memory | Processed in their cloud |
Frequently asked questions
Not a full one, and we would rather say so up front. It replaces one mechanic: cutting a recording by deleting words, plus removal of the English filler sounds. If your Descript workflow involves multi-track video, AI voice, screen recording or shared projects with comments, nothing here covers that.
Nothing. No account, no credit balance, no watermark and no paid tier of the editor. The work runs on your own machine, so there is no per-minute cost for us to pass on to you.
No. Open the page and drop in a file. Descript asks you to create an account before the first edit; this does not.
No. The browser reads the file off your disk and the speech model runs on your own hardware, so the media and the text never leave the device. Descript processes your media in its cloud. The only network requests here are the app itself, the one-time model download, and anonymous product analytics with no file data in them.
Honestly, it is worse. Only small models fit in a browser tab, so accuracy drops on accents, crosstalk, background noise and non-English audio. If the words have to be right, run the file through ConvertAudioToText and bring the result into the editor to do the cutting.
Everything is held in memory, so a laptop is comfortable with clips up to roughly ten to twenty minutes and gets slow or runs out past that. Longer recordings are exactly the case for running the file through ConvertAudioToText first.
A recent Chromium desktop browser, so Chrome, Edge, Brave or Arc. Firefox works but is slower. Safari and most mobile browsers are not reliable for this workload. Descript's desktop app has no such constraint.
They are different products. The editor is free and runs locally. ConvertAudioToText is the hosted service for long files, 99+ languages, speaker separation and publish-grade accuracy, and it has its own side-by-side comparison with Descript.