Make Autocut
Try free
By Yohann Kipfer · Last updated July 2026

A Cleanvoice alternative for people who edit

Cleanvoice, Descript and Auphonic decide for you and hand back a rendered file. Make Autocut does one thing instead: it finds the silence, shows you exactly what it’s about to cut, and gives you an editable timeline — Premiere, Final Cut, Resolve, or anything via OpenTimelineIO. You keep the edit. Nothing is a black box.

In short: if you need filler words and breaths gone, use Cleanvoice or Descript — we don’t do that. If you want the dead air cut and the cut handed back to your editor to finish by hand, cheaper and in the browser, that’s us. Free tier: 5 min, 200 MB, one file a day, no account.

Why editors go looking for a Cleanvoice alternative

Cleanvoice is good software. People still leave, and it’s usually one of four honest reasons — not because the AI is bad, but because of what the whole category is shaped like.

You want the cut, not a finished file

The cloud tools return a rendered export. If you edit for a living, that's the wrong artefact — you want the decisions back as a timeline you can slide and undo, not a flattened mp3 you have to trust.

You don't want another black box

You press a button, something happens, you hope it was right. There's no threshold to see, no waveform to check, no way to say “that pause was carrying weight, leave it.” Control is the thing that's missing.

The price adds up for one feature

A monthly subscription or a $129 desktop app is a lot when the job you actually do is “take the silence out.” The maths gets worse if you only cut a few files a month.

It's locked to one editor, or to none

Some tools export to a single NLE; the cloud ones export to none. The day the same cut has to open in Premiere and in Resolve, a per-app format is a dead end.

Ten silence tools, side by side

Everyone’s comparison table leaves off the column that would lose it a checkmark. Here’s the one we win and most tables skip: editable timeline export, and specifically whether it’s locked to one editor or universal. Recut and Timebolt do export timelines — that’s their strength — but to per-app formats. Only one row below carries OpenTimelineIO.

Silence-removal tools compared by what they remove, timeline export, where they run and price
ToolWhat it removesEditable timeline exportRunsPrice (published, changes)
Make AutocutSilence onlyPremiere XML, FCPXML, EDL + OTIO (universal)Browser, no installFree 5 min/day · Pro from a few $ /mo
CleanvoiceSilence, filler words, stutters, mouth soundsNo — returns a processed fileCloudSubscription / hour-based credits, ~$10/mo up
DescriptFiller words + silence, via transcriptSome, via its own export to a couple of appsHeavy desktop appFree tier, then ~$24–30/mo
AuphonicNoise, loudness, filler (beta) — not a cutterNo — audio + multitrack outCloud~2 free hours/mo, then paid by hour
RecutSilence (and by transcript)Yes — per-app XML (Premiere, FCP, Resolve, Audition)Desktop appOne-time, around $129
TimeboltSilence + manual jump-cutsYes — per-app XML + EDLDesktop appOne-time tiers, ~$50–347
GlingSilence + filler (YouTube-first)Yes — per-app XMLDesktop / cloudSubscription
VEEDSilence + a whole online editorNo — renders in-appCloudFree tier, then subscription
AudacitySilence (Truncate Silence), by handNo — audio onlyDesktop, freeFree (open source)
KapwingSilence + online editorNo — renders in-appCloudFree tier, then subscription

Prices are what each vendor published around July 2026 and are starting points, not quotes — tiers move faster than this page. The one claim that stays true is the third column: nobody else on this list exports OpenTimelineIO, so nobody else gives you a cut that opens in any editor. Read the “what it removes” column before the price one — half these tools do more than we do, on purpose.

The one thing no black box gives you: your edit back

A cloud silence remover returns a file. That’s fine until you’re an editor and the file is wrong in a way only you can hear — a pause that was doing work, a join that clips a breath. Now you’re fighting a rendered artefact instead of moving a clip edge. Handing back a timeline flips that: every cut is a straight cut you can slide, ripple or delete in the tool you already use.

Recut and Timebolt understand this and export to editors too — credit where it’s due. Where we go further is the format. They write per-app XML: pick your editor, get its dialect. We write those and OpenTimelineIO, the open interchange the film industry built precisely so a timeline isn’t hostage to one vendor. It’s pure string templating over the cut list we already have in the database — no ffmpeg, no re-encode, your media never opened. Here’s the shape of it, from src/lib/timeline.ts:

{
  "OTIO_SCHEMA": "Timeline.1",
  "tracks": { "OTIO_SCHEMA": "Stack.1", "children": [
    { "OTIO_SCHEMA": "Track.1", "kind": "Video", "children": [
      { "OTIO_SCHEMA": "Clip.1", "name": "interview.mp4 #1",
        "source_range": { "start_time": {...}, "duration": {...} },
        "media_reference": { "target_url": "file://localhost/interview.mp4" } }
      // one clip per kept segment, silences rippled out
    ] }
  ] }
}

Comparing Descript specifically? The Descript alternative page goes deeper on the subscription-and-install trade and the cases where Descript still wins.

The four editor pages document the round trip end to end: Premiere Pro, Final Cut Pro and DaVinci Resolve — import steps, the frame-rate trap, and why every clip relinks.

Read this before you switch: what we don’t do

A comparison written by the tool always has the empty boxes on the other side. Ours don’t. Here is everything Cleanvoice and friends do that we flatly don’t — and if one of these is your real problem, we’re the wrong page and one of them is the right one.

Filler words, stutters, breaths

We cut silence, full stop. An “um” is sound above the noise floor, so a loudness detector keeps it. Removing verbal tics needs transcription — Cleanvoice, Descript and Podcastle do it, we don't. If your recording is clean but you say “like” a lot, don't come here.

Noise reduction and loudness

No denoise, no speech enhancement, no loudness targets. Auphonic and Adobe Podcast own that half of post. Clean the audio there first, then cut the silence here — the order matters and they stack.

Transcription and text-based editing

You can't edit by editing a transcript. There's no transcript at all — detection is volume-based. That's why it works in any language, and why it can't do what Descript or Riverside do.

Multitrack and multi-cam

One mixed-down file in, one timeline out. No synced multi-mic, no multi-camera. Recut and Auphonic multitrack handle that; a two-mic podcast where you need per-mic control is not our case.

Batch processing

One file at a time. Timebolt and Auphonic batch a folder; we don't yet. For a back catalogue of 40 episodes overnight, that's a real gap.

The trade is deliberate. We do one job and do it transparently, for a free tier and a couple of dollars, instead of half a dozen jobs behind a subscription. If that one job is your bottleneck, that’s a good deal. If it isn’t, we just told you so.

The straight trade, tool by tool

Named comparisons, with the real reason to pick them over us on top and the real reason to pick us underneath. Prices are as published around July 2026.

vs Cleanvoice

Cloud, ~$10/mo and up (hour-based credits)

Pick them if: It removes what we can't: filler words, stutters, lip smacks, breaths. If your problem is “um” every third sentence, Cleanvoice is the right tool and this page isn't.

Pick us if: You never see how Cleanvoice decided. We show the waveform, the exact dB threshold we picked from your file, and let you drag it before anything is cut — then you leave with an editable timeline, not a rendered mp3.

vs Descript

Free tier, then roughly $24–30/mo

Pick them if: Descript edits video by editing a transcript, kills filler words, and is a genuinely good all-in-one. If you want to write your edit, use Descript.

Pick us if: It's a heavy app and a subscription for a whole workflow. If all you actually need is the silence gone and the cut handed back to Premiere or Resolve, you're renting a suite to use one feature — and Descript's own NLE export is narrower than four formats plus OTIO.

vs Auphonic

~2 free hours/month, then paid by processing hour

Pick them if: Auphonic is loudness normalisation, noise reduction and multitrack levelling — the audio-mastering half of post. It is excellent at that and we do none of it.

Pick us if: Auphonic isn't a silence cutter and doesn't hand you a timeline. Run Auphonic to clean the audio, then run us to cut the dead air and export the edit. They stack; they don't compete.

vs Recut

One-time, around $129

Pick them if: Recut lives on your desktop, handles multi-cam and multi-mic, and exports straight into Premiere, Final Cut, Resolve and Audition. For a multitrack interview it does things we simply can't — we're mono-file.

Pick us if: $129 up front versus a free file today, in the browser, on any machine you can't install software on. And Recut stops at per-app XML; we add OpenTimelineIO, so the same cut opens in an editor Recut never wrote a serialiser for.

vs Timebolt

One-time tiers, roughly $50–347

Pick them if: Timebolt is fast at manual jump-cutting on top of silence detection, and it batches. If you cut the same way every day and want keyboard-driven speed, it earns its place.

Pick us if: The top Timebolt tier is north of $300. Ours is a free tier that actually works plus cheap Pro, and the same universal-export argument: per-app XML from them, per-app XML plus OTIO from us.

How it works, step by step

Four steps, and the difference from a black box is step three: you get to see and change every decision before a single frame is cut.

  1. 1. Drop the file in — no account for the first one

    Upload the audio or video you're going to cut, up to 5 minutes and 200 MB on the free tier. It goes straight to storage; nobody watches it and it's deleted an hour later. Give us the real file, not a bounced MP3, if you want a video timeline back.

  2. 2. We measure your noise floor, not a global default

    A fixed threshold is the wrong idea — a treated booth and a laptop with a fan running don't share a noise floor, so −40 dB means something different on each. We run volumedetect, read your file's mean volume, and set the threshold to mean − 5 dB (clamped to −50…−25, falling back to -35 dB only if the probe returns nothing). You see the number it chose.

  3. 3. Tune the threshold, minimum silence and padding — live

    This is the part the black boxes hide. Drag the dB threshold and the waveform re-highlights what's about to go. Raise the minimum-silence length (default 0.8 s) to keep meaningful teaching pauses. Change the padding (default 150 ms each end, inside the 0.1–0.3 s range editors expect) so joins don't sound clipped. A 1.00 s pause at defaults loses 0.70 s and keeps 0.30 s of air.

  4. 4. Take the render, or take the timeline

    Export the cut MP4/MP3 if you just want the clean file. Or click your editor and download an FCPXML, Premiere XML, EDL or OTIO — the cut list as an editable timeline, no render, no re-encode. Every clip lands offline and asks to relink once (a browser upload strips the disk path, so we never knew where your media lives; an honest relink beats a silent mislink).

Every dB and default above is the real value the code uses, not a marketing round number. The main auto cut audio page has 155 threshold measurements on five public-domain recordings you can download and re-run — including the one where the tool honestly fails to cut anything.

Which editor do you use? Pick the export

The naming is a trap: FCPXML and FCP7 XML share three letters and nothing else. Send the wrong dialect and the import silently does nothing. So you don’t choose a format — you choose your editor, and we write the right one.

Which timeline export to pick for each editor
Your editorExport to pickWhy
Premiere ProPremiere (FCP7 XML)Premiere reads the old xmeml dialect, not FCPXML — despite the shared name they're unrelated formats.
Final Cut ProFCPXMLIts own native format. Arrives as a library/event/project you import directly, from Final Cut 10.4.9 on.
DaVinci ResolveFCPXML (or EDL)Resolve reads FCPXML as a well-behaved guest; EDL is the fallback when you want the plainest possible interchange.
Avid, anything else, or codeOTIOOpenTimelineIO is the editor-agnostic one. This is the row no other cutter has — one cut list, any NLE.

OpenTimelineIO is documented by the Academy Software Foundation at opentimelineio.readthedocs.io. It’s the reason the same cut can leave here and open in an editor we’ve never tested.

Cut a file and see the difference

First one is free, no account. You see the waveform, the pauses and the threshold it picked before anything is cut — then take the render, or take the timeline. No black box.

Drop your audio or video file
or browse
MP3, WAV, M4A, MP4, MOV, MKV, WEBM — up to 200 MB

Want the exact import steps for your editor? The Premiere, Final Cut and Resolve pages walk the round trip click by click.

Questions people actually ask

Is Make Autocut really a Cleanvoice alternative if it doesn't remove filler words?

Only for one job: cutting silence. If you want “um”, “uh”, stutters and mouth noises gone, Cleanvoice and Descript do that with transcription and we honestly don't — an “um” is sound, not silence, and a loudness detector can't see it. Where we're the better alternative is when you don't want a black box: you want to see the pauses, control the threshold, and leave with an editable timeline instead of a rendered file. Pick the tool by which of those two problems you actually have.

Does it keep the background music or cut into it?

It cuts on loudness, so anything above the noise floor is “not silence” and survives — including music. That's usually what you want on a podcast bed. The failure case is the opposite: on a noisy source where the floor sits close to the voice, it removes almost nothing. Our public bench includes an Apollo 11 comms tape (constant hiss) where the tool cuts 0.0% at auto, and we left that in the data rather than hide it.

How much time does it actually save?

Not by removing seconds — that's small. On our bench (2026-07-16) a prepared voice-over gave back 3.3% of its length and a recorded lecture 1.4%. The real saving is the razor cuts you don't place by hand: that voice-over needed 16 cuts in five minutes, the lecture 7. Scale a lecture to 45 minutes and that's roughly 63 manual cuts you'd otherwise make one at a time. We do those in one click and hand you the timeline to fix by ear.

Which spoken languages does it work in?

All of them. Detection is volume-based, not transcription-based, so English, French, Japanese, a whistle or a dog barking are treated the same — sound above the floor is kept, silence below it is cut. That's the flip side of not removing filler words: nothing here depends on understanding the words, so nothing here is limited to the languages a speech model was trained on.

Is the audio quality degraded?

On the timeline export, no — nothing is re-encoded. The timeline endpoint reads the cut list from the database and templates a string; ffmpeg never opens your media. Your editor plays your original file off your own disk after you relink. The MP4/MP3 export does re-encode (H.264 CRF 20, or LAME -q:a 2), so if preserving the master matters, take the timeline and never press the render button.

Can I undo a cut, or is it destructive like a rendered file?

Your source is never touched — we work on an upload and delete it an hour later. Take the timeline export and every cut is fully reversible in your NLE: it's a straight cut you can slide, ripple back, or delete. That's the whole argument for a timeline over a flattened file. A black box hands you an mp3 and you live with its decisions; we hand you an edit you can disagree with, cut by cut.

Which export do I pick for my editor?

Premiere Pro reads FCP7 XML — click “Premiere”. Final Cut Pro reads FCPXML. DaVinci Resolve reads FCPXML too, or EDL if you'd rather. Anything else — Avid, a scripted pipeline, an editor we've never heard of — takes OpenTimelineIO (OTIO), the universal one. When in doubt, OTIO opens the widest. Send the wrong dialect and the import silently does nothing, which is exactly why we label the buttons by editor, not by format.

What does it cost, and is the free tier real?

The free tier is one file a day, up to 5 minutes and 200 MB, no account, no watermark, all export formats unlocked. It's genuinely free, not a trial that expires. The honest limit: 5 minutes won't hold a real podcast episode, so if this becomes part of how you edit you'll move to Pro (60 min, 2 GB, no daily cap) — which is still a fraction of a $129 desktop app or a $24-a-month suite.

Try it on a real file. No account, no card.

5 minutes, free, no watermark — and every export format, timeline included, is unlocked.

Cut my file free