Descript

descript.comcontributed by Samuele Ongaro

KEEP IT

The value is the network, the data or the infrastructure. Keep paying.

Edits audio and video by editing text: it transcribes the recording, and deleting a word deletes that piece of the timeline. It adds filler removal, sound cleanup and synthetic voice on top.

Promptfree, for everyone, and the only version there is
Do not build this. Editing video by editing its transcript sounds simple and is not.

Every word must map to a frame range, the timeline must stay consistent as text is deleted, and cuts must land on sensible boundaries rather than mid-syllable. Then filler removal that does not leave audible gaps, sound cleanup that works on ordinary recordings, and synthetic voice for a word you got wrong. That combination is the product.

WHAT TO DO INSTEAD
- Keep it if you edit spoken video regularly. It genuinely removes the timeline, which is the barrier for everybody who is not an editor
- If you record occasionally, the clip-based approach — record a take, keep or discard, re-record a bad one — gets most of the benefit and is covered under Tella elsewhere in this catalogue

WHAT YOU CAN BUILD, AND SHOULD

**Transcript-driven rough cuts**, which is the ninety per cent that is not hard:

STACK
- whisper.cpp or an equivalent local model with word-level timestamps, and ffmpeg

WHAT TO IMPLEMENT
- Transcribe with per-word timings, stored as rows rather than a blob
- Silence detection, and a list of every filler word with its exact span
- Generate an edit decision list: the segments to keep, expressed as timestamps
- Apply it with ffmpeg in one pass, cutting on the nearest zero crossing so a cut does not click
- A short crossfade at each join — a few tens of milliseconds — which is what makes filler removal inaudible. Hard cuts are why homemade versions sound wrong
- Review before applying, always. An automatic cut that removed half a sentence is discovered by the audience

AND THE FREE WIN
Most of what makes a recording sound bad is not editing. Loudness normalisation to a broadcast target, a high-pass filter to remove rumble, and gentle noise reduction — three ffmpeg filters — improve a recording more than any amount of cutting.

WHAT NOT TO ATTEMPT
Synthetic voice cloning of yourself. The models exist; the ethical and legal ground around a voice that can say anything is not somewhere to wander casually.

THE ONE-LINE VERSION
Buy the transcript editor. Build the rough cut and the three audio filters, which is most of the improvement for none of the difficulty.

What you lose

  • Editing video by editing its transcript, which is the whole idea and is genuinely hard to build
  • Filler-word removal and studio sound that work well enough to use unattended
  • Overdub and voice cloning with the consent machinery around them
  • Multi-track timeline editing with scenes, layers and transitions
  • Rendering and export queues that do not run on your own laptop

If you would rather not build

  • Whisper plus ffmpeg — the do-it-yourself path, no interface
  • Riverside — paid, recording-first with transcript editing

What it costs

read from their page 15 Aug 2026

PlanBilled monthlyBilled yearlyLast read
—$24/mo—15 Aug 2026

Their pricing page is where these came from. Seeing a different price? Tell us.

The escape hatch

open source · no votes, no paid placement

whisper.cpp

$0

Runs speech-to-text locally on ordinary hardware, with word-level timestamps.

ggml-org/whisper.cppfree · open source

Kdenlive

$0

Full open-source video editor, for when the timeline is what you actually need.

KDE/kdenlivefree · open source

Why this verdict

our own opinion · changed only by a person

29/100

Verdict no at 29: transcript-driven trimming is buildable and worth having, and the prompt covers the two traps — word-level timings and crossfading joins. Everything else Descript does, from studio sound to voice, is research work, not a weekend.

History

tracked since 9 Aug 2026 · nothing is ever overwritten

Interest · last 30 dayspeak 2/day
views01230 Aug4 Sept9 Sept14 Sept19 Sept24 Sept28 Sept
— views— prompt copies none yet— votes none yet

Questions about Descript

answered from the record above

Is Descript free?

No — the plan we track is $24 a month. Hobbyist around $24/month billed monthly, $12 annually; transcription hours and export resolution are the limits that bite.

Can you replace Descript by building your own?

KEEP IT. The value is the network, the data or the infrastructure. Keep paying. Replacement score 29 out of 100, build time longer than it saves. Read what you lose before you decide.

How much does Descript cost?

$24 a month on Hobbyist — $288 a year. Recorded 9 Aug 2026.

What do you lose by replacing Descript?

Editing video by editing its transcript, which is the whole idea and is genuinely hard to build; Filler-word removal and studio sound that work well enough to use unattended; Overdub and voice cloning with the consent machinery around them; Multi-track timeline editing with scenes, layers and transitions; Rendering and export queues that do not run on your own laptop. If any of those carry weight for you, keep paying.

Is there an open-source alternative to Descript?

Yes: whisper.cpp, Kdenlive. The prompt on this page is for when you want it your way instead.

Related entries

same category first, most replaced first

All 26 in Screen recording & video

Not sending yet

Every week, something stops being worth paying for.

New verdicts, prices that moved, entries added. One email a week. Unsubscribe in one click. Nothing is being sent yet — your address is kept here, and the first issue is the first thing it is used for.

free forever · no tracking pixel · stored here, never passed to anyone

Esc