A Cleanvoice alternative for people who edit
Cleanvoice, Descript and Auphonic decide for you and hand back a rendered file. Make Autocut does one thing instead: it finds the silence, shows you exactly what it’s about to cut, and gives you an editable timeline — Premiere, Final Cut, Resolve, or anything via OpenTimelineIO. You keep the edit. Nothing is a black box.
In short: if you need filler words and breaths gone, use Cleanvoice or Descript — we don’t do that. If you want the dead air cut and the cut handed back to your editor to finish by hand, cheaper and in the browser, that’s us. Free tier: 5 min, 200 MB, one file a day, no account.
Why editors go looking for a Cleanvoice alternative
Cleanvoice is good software. People still leave, and it’s usually one of four honest reasons — not because the AI is bad, but because of what the whole category is shaped like.
You want the cut, not a finished file
The cloud tools return a rendered export. If you edit for a living, that's the wrong artefact — you want the decisions back as a timeline you can slide and undo, not a flattened mp3 you have to trust.
You don't want another black box
You press a button, something happens, you hope it was right. There's no threshold to see, no waveform to check, no way to say “that pause was carrying weight, leave it.” Control is the thing that's missing.
The price adds up for one feature
A monthly subscription or a $129 desktop app is a lot when the job you actually do is “take the silence out.” The maths gets worse if you only cut a few files a month.
It's locked to one editor, or to none
Some tools export to a single NLE; the cloud ones export to none. The day the same cut has to open in Premiere and in Resolve, a per-app format is a dead end.
Ten silence tools, side by side
Everyone’s comparison table leaves off the column that would lose it a checkmark. Here’s the one we win and most tables skip: editable timeline export, and specifically whether it’s locked to one editor or universal. Recut and Timebolt do export timelines — that’s their strength — but to per-app formats. Only one row below carries OpenTimelineIO.
| Tool | What it removes | Editable timeline export | Runs | Price (published, changes) |
|---|---|---|---|---|
| Make Autocut | Silence only | Premiere XML, FCPXML, EDL + OTIO (universal) | Browser, no install | Free 5 min/day · Pro from a few $ /mo |
| Cleanvoice | Silence, filler words, stutters, mouth sounds | No — returns a processed file | Cloud | Subscription / hour-based credits, ~$10/mo up |
| Descript | Filler words + silence, via transcript | Some, via its own export to a couple of apps | Heavy desktop app | Free tier, then ~$24–30/mo |
| Auphonic | Noise, loudness, filler (beta) — not a cutter | No — audio + multitrack out | Cloud | ~2 free hours/mo, then paid by hour |
| Recut | Silence (and by transcript) | Yes — per-app XML (Premiere, FCP, Resolve, Audition) | Desktop app | One-time, around $129 |
| Timebolt | Silence + manual jump-cuts | Yes — per-app XML + EDL | Desktop app | One-time tiers, ~$50–347 |
| Gling | Silence + filler (YouTube-first) | Yes — per-app XML | Desktop / cloud | Subscription |
| VEED | Silence + a whole online editor | No — renders in-app | Cloud | Free tier, then subscription |
| Audacity | Silence (Truncate Silence), by hand | No — audio only | Desktop, free | Free (open source) |
| Kapwing | Silence + online editor | No — renders in-app | Cloud | Free tier, then subscription |
Prices are what each vendor published around July 2026 and are starting points, not quotes — tiers move faster than this page. The one claim that stays true is the third column: nobody else on this list exports OpenTimelineIO, so nobody else gives you a cut that opens in any editor. Read the “what it removes” column before the price one — half these tools do more than we do, on purpose.
The one thing no black box gives you: your edit back
A cloud silence remover returns a file. That’s fine until you’re an editor and the file is wrong in a way only you can hear — a pause that was doing work, a join that clips a breath. Now you’re fighting a rendered artefact instead of moving a clip edge. Handing back a timeline flips that: every cut is a straight cut you can slide, ripple or delete in the tool you already use.
Recut and Timebolt understand this and export to editors too — credit where it’s due. Where we go further is the format. They write per-app XML: pick your editor, get its dialect. We write those and OpenTimelineIO, the open interchange the film industry built precisely so a timeline isn’t hostage to one vendor. It’s pure string templating over the cut list we already have in the database — no ffmpeg, no re-encode, your media never opened. Here’s the shape of it, from src/lib/timeline.ts:
{
"OTIO_SCHEMA": "Timeline.1",
"tracks": { "OTIO_SCHEMA": "Stack.1", "children": [
{ "OTIO_SCHEMA": "Track.1", "kind": "Video", "children": [
{ "OTIO_SCHEMA": "Clip.1", "name": "interview.mp4 #1",
"source_range": { "start_time": {...}, "duration": {...} },
"media_reference": { "target_url": "file://localhost/interview.mp4" } }
// one clip per kept segment, silences rippled out
] }
] }
}Comparing Descript specifically? The Descript alternative page goes deeper on the subscription-and-install trade and the cases where Descript still wins.
The four editor pages document the round trip end to end: Premiere Pro, Final Cut Pro and DaVinci Resolve — import steps, the frame-rate trap, and why every clip relinks.
Read this before you switch: what we don’t do
A comparison written by the tool always has the empty boxes on the other side. Ours don’t. Here is everything Cleanvoice and friends do that we flatly don’t — and if one of these is your real problem, we’re the wrong page and one of them is the right one.
Filler words, stutters, breaths
We cut silence, full stop. An “um” is sound above the noise floor, so a loudness detector keeps it. Removing verbal tics needs transcription — Cleanvoice, Descript and Podcastle do it, we don't. If your recording is clean but you say “like” a lot, don't come here.
Noise reduction and loudness
No denoise, no speech enhancement, no loudness targets. Auphonic and Adobe Podcast own that half of post. Clean the audio there first, then cut the silence here — the order matters and they stack.
Transcription and text-based editing
You can't edit by editing a transcript. There's no transcript at all — detection is volume-based. That's why it works in any language, and why it can't do what Descript or Riverside do.
Multitrack and multi-cam
One mixed-down file in, one timeline out. No synced multi-mic, no multi-camera. Recut and Auphonic multitrack handle that; a two-mic podcast where you need per-mic control is not our case.
Batch processing
One file at a time. Timebolt and Auphonic batch a folder; we don't yet. For a back catalogue of 40 episodes overnight, that's a real gap.
The trade is deliberate. We do one job and do it transparently, for a free tier and a couple of dollars, instead of half a dozen jobs behind a subscription. If that one job is your bottleneck, that’s a good deal. If it isn’t, we just told you so.
The straight trade, tool by tool
Named comparisons, with the real reason to pick them over us on top and the real reason to pick us underneath. Prices are as published around July 2026.
vs Cleanvoice
Cloud, ~$10/mo and up (hour-based credits)Pick them if: It removes what we can't: filler words, stutters, lip smacks, breaths. If your problem is “um” every third sentence, Cleanvoice is the right tool and this page isn't.
Pick us if: You never see how Cleanvoice decided. We show the waveform, the exact dB threshold we picked from your file, and let you drag it before anything is cut — then you leave with an editable timeline, not a rendered mp3.
vs Descript
Free tier, then roughly $24–30/moPick them if: Descript edits video by editing a transcript, kills filler words, and is a genuinely good all-in-one. If you want to write your edit, use Descript.
Pick us if: It's a heavy app and a subscription for a whole workflow. If all you actually need is the silence gone and the cut handed back to Premiere or Resolve, you're renting a suite to use one feature — and Descript's own NLE export is narrower than four formats plus OTIO.
vs Auphonic
~2 free hours/month, then paid by processing hourPick them if: Auphonic is loudness normalisation, noise reduction and multitrack levelling — the audio-mastering half of post. It is excellent at that and we do none of it.
Pick us if: Auphonic isn't a silence cutter and doesn't hand you a timeline. Run Auphonic to clean the audio, then run us to cut the dead air and export the edit. They stack; they don't compete.
vs Recut
One-time, around $129Pick them if: Recut lives on your desktop, handles multi-cam and multi-mic, and exports straight into Premiere, Final Cut, Resolve and Audition. For a multitrack interview it does things we simply can't — we're mono-file.
Pick us if: $129 up front versus a free file today, in the browser, on any machine you can't install software on. And Recut stops at per-app XML; we add OpenTimelineIO, so the same cut opens in an editor Recut never wrote a serialiser for.
vs Timebolt
One-time tiers, roughly $50–347Pick them if: Timebolt is fast at manual jump-cutting on top of silence detection, and it batches. If you cut the same way every day and want keyboard-driven speed, it earns its place.
Pick us if: The top Timebolt tier is north of $300. Ours is a free tier that actually works plus cheap Pro, and the same universal-export argument: per-app XML from them, per-app XML plus OTIO from us.
How it works, step by step
Four steps, and the difference from a black box is step three: you get to see and change every decision before a single frame is cut.
1. Drop the file in — no account for the first one
Upload the audio or video you're going to cut, up to 5 minutes and 200 MB on the free tier. It goes straight to storage; nobody watches it and it's deleted an hour later. Give us the real file, not a bounced MP3, if you want a video timeline back.
2. We measure your noise floor, not a global default
A fixed threshold is the wrong idea — a treated booth and a laptop with a fan running don't share a noise floor, so −40 dB means something different on each. We run volumedetect, read your file's mean volume, and set the threshold to mean − 5 dB (clamped to −50…−25, falling back to -35 dB only if the probe returns nothing). You see the number it chose.
3. Tune the threshold, minimum silence and padding — live
This is the part the black boxes hide. Drag the dB threshold and the waveform re-highlights what's about to go. Raise the minimum-silence length (default 0.8 s) to keep meaningful teaching pauses. Change the padding (default 150 ms each end, inside the 0.1–0.3 s range editors expect) so joins don't sound clipped. A 1.00 s pause at defaults loses 0.70 s and keeps 0.30 s of air.
4. Take the render, or take the timeline
Export the cut MP4/MP3 if you just want the clean file. Or click your editor and download an FCPXML, Premiere XML, EDL or OTIO — the cut list as an editable timeline, no render, no re-encode. Every clip lands offline and asks to relink once (a browser upload strips the disk path, so we never knew where your media lives; an honest relink beats a silent mislink).
Every dB and default above is the real value the code uses, not a marketing round number. The main auto cut audio page has 155 threshold measurements on five public-domain recordings you can download and re-run — including the one where the tool honestly fails to cut anything.
Which editor do you use? Pick the export
The naming is a trap: FCPXML and FCP7 XML share three letters and nothing else. Send the wrong dialect and the import silently does nothing. So you don’t choose a format — you choose your editor, and we write the right one.
| Your editor | Export to pick | Why |
|---|---|---|
| Premiere Pro | Premiere (FCP7 XML) | Premiere reads the old xmeml dialect, not FCPXML — despite the shared name they're unrelated formats. |
| Final Cut Pro | FCPXML | Its own native format. Arrives as a library/event/project you import directly, from Final Cut 10.4.9 on. |
| DaVinci Resolve | FCPXML (or EDL) | Resolve reads FCPXML as a well-behaved guest; EDL is the fallback when you want the plainest possible interchange. |
| Avid, anything else, or code | OTIO | OpenTimelineIO is the editor-agnostic one. This is the row no other cutter has — one cut list, any NLE. |
OpenTimelineIO is documented by the Academy Software Foundation at opentimelineio.readthedocs.io. It’s the reason the same cut can leave here and open in an editor we’ve never tested.
Cut a file and see the difference
First one is free, no account. You see the waveform, the pauses and the threshold it picked before anything is cut — then take the render, or take the timeline. No black box.
Want the exact import steps for your editor? The Premiere, Final Cut and Resolve pages walk the round trip click by click.
Questions people actually ask
Is Make Autocut really a Cleanvoice alternative if it doesn't remove filler words?
▾
Only for one job: cutting silence. If you want “um”, “uh”, stutters and mouth noises gone, Cleanvoice and Descript do that with transcription and we honestly don't — an “um” is sound, not silence, and a loudness detector can't see it. Where we're the better alternative is when you don't want a black box: you want to see the pauses, control the threshold, and leave with an editable timeline instead of a rendered file. Pick the tool by which of those two problems you actually have.
Does it keep the background music or cut into it?
▾
It cuts on loudness, so anything above the noise floor is “not silence” and survives — including music. That's usually what you want on a podcast bed. The failure case is the opposite: on a noisy source where the floor sits close to the voice, it removes almost nothing. Our public bench includes an Apollo 11 comms tape (constant hiss) where the tool cuts 0.0% at auto, and we left that in the data rather than hide it.
How much time does it actually save?
▾
Not by removing seconds — that's small. On our bench (2026-07-16) a prepared voice-over gave back 3.3% of its length and a recorded lecture 1.4%. The real saving is the razor cuts you don't place by hand: that voice-over needed 16 cuts in five minutes, the lecture 7. Scale a lecture to 45 minutes and that's roughly 63 manual cuts you'd otherwise make one at a time. We do those in one click and hand you the timeline to fix by ear.
Which spoken languages does it work in?
▾
All of them. Detection is volume-based, not transcription-based, so English, French, Japanese, a whistle or a dog barking are treated the same — sound above the floor is kept, silence below it is cut. That's the flip side of not removing filler words: nothing here depends on understanding the words, so nothing here is limited to the languages a speech model was trained on.
Is the audio quality degraded?
▾
On the timeline export, no — nothing is re-encoded. The timeline endpoint reads the cut list from the database and templates a string; ffmpeg never opens your media. Your editor plays your original file off your own disk after you relink. The MP4/MP3 export does re-encode (H.264 CRF 20, or LAME -q:a 2), so if preserving the master matters, take the timeline and never press the render button.
Can I undo a cut, or is it destructive like a rendered file?
▾
Your source is never touched — we work on an upload and delete it an hour later. Take the timeline export and every cut is fully reversible in your NLE: it's a straight cut you can slide, ripple back, or delete. That's the whole argument for a timeline over a flattened file. A black box hands you an mp3 and you live with its decisions; we hand you an edit you can disagree with, cut by cut.
Which export do I pick for my editor?
▾
Premiere Pro reads FCP7 XML — click “Premiere”. Final Cut Pro reads FCPXML. DaVinci Resolve reads FCPXML too, or EDL if you'd rather. Anything else — Avid, a scripted pipeline, an editor we've never heard of — takes OpenTimelineIO (OTIO), the universal one. When in doubt, OTIO opens the widest. Send the wrong dialect and the import silently does nothing, which is exactly why we label the buttons by editor, not by format.
What does it cost, and is the free tier real?
▾
The free tier is one file a day, up to 5 minutes and 200 MB, no account, no watermark, all export formats unlocked. It's genuinely free, not a trial that expires. The honest limit: 5 minutes won't hold a real podcast episode, so if this becomes part of how you edit you'll move to Pro (60 min, 2 GB, no daily cap) — which is still a fraction of a $129 desktop app or a $24-a-month suite.
Try it on a real file. No account, no card.
5 minutes, free, no watermark — and every export format, timeline included, is unlocked.
Cut my file free