Transcribes meetings and conversations: joins calls, separates speakers, produces a searchable transcript, and summarises the outcome.
Build me meeting transcription that replaces Otter — and read the two obstacles first. **A bot that joins the call** so nobody has to remember to record is the feature, and building one means driving a meeting client programmatically, which each platform makes awkward in its own way. **Speaker separation with overlapping voices** is the other, and it is a research problem where local models are noticeably weaker. And before either: in many places recording a conversation requires everyone's consent, so the announcement is not a courtesy, it is the law. STACK - Node 20+ with Fastify - SQLite through better-sqlite3, WAL mode, with FTS5 - whisper.cpp or an equivalent local model for transcription, and a diarisation model for speakers - ffmpeg for audio handling - Caddy in front CONSENT, WHICH COMES FIRST - Announce the recording at the start, audibly and visibly, before anything is captured - Record that the announcement happened, with a timestamp, per meeting - A way for anybody to object and stop it. Not a setting somebody has to find afterwards - Write the jurisdictions you have checked in the README, and take advice. Recording a colleague without telling them is not a technical decision THE DATA MODEL - meetings: id, title, source, external_id, starts_at_utc, ends_at_utc, timezone, audio_path, duration_ms, status, consent_announced_at, created_by - participants: id, meeting_id, name, email, speaker_label, joined_at, left_at - segments: id, meeting_id, speaker_label, start_ms, end_ms, text, confidence, is_edited — the transcript, as segments rather than one blob, which is what makes search, editing and playback alignment possible - summaries: id, meeting_id, kind, body_md, model_version, generated_at, is_edited - actions: id, meeting_id, segment_id, text, assignee, due_at, status — extracted, then confirmed by a human - highlights, comments, shares - Audio kept as the original, with the transcript derived. Never discard the audio while the transcript is the only record, because transcription is imperfect and the recording is the evidence GETTING THE AUDIO IN - The easy path, and the one to build first: upload a recording, or point at a file the meeting platform already produced. That covers most of the value with none of the difficulty - A browser recorder for an in-person conversation, capturing the microphone directly - The bot that joins a call is the hard path: a headless client, an account, and per-platform behaviour that changes. Some platforms offer a proper API for this; check before writing anything, because if yours does, the problem disappears - Live transcription during a call is a streaming model with partial results, which is a different pipeline from batch. Build batch first TRANSCRIPTION - Chunk the audio with overlap so a word split across a boundary is not lost, and stitch the results - Store confidence per segment, and show low-confidence text differently. A transcript that hides its uncertainty gets quoted wrongly - A custom vocabulary of names, products and jargon, applied as a post-pass — this is the single biggest quality improvement available and it costs almost nothing - Speaker diarisation produces labels, not names. Let a human name each speaker once per meeting and remember the voice for next time, which is enough without solving the general problem THE TRANSCRIPT - Editable, with edits marked and the original kept - Click a line to play from that moment. That alignment is what makes a transcript usable rather than just readable - Search across every meeting with FTS5, with the matching line and the timestamp - Export as text, Markdown, subtitles and a document SUMMARIES, HONESTLY - Generate from the transcript with a local model if you have one, and label it as generated - Never present a summary without the transcript beside it. A summary is a lossy reading of a conversation people are accountable for, and the source must be one click away - Action items extracted and then confirmed by a person before they become tasks. An automatically assigned action nobody agreed to is worse than none PRIVACY - This holds recordings of colleagues talking. Encrypt at rest, restrict access, log every playback, and set a real retention with a sweeper that deletes - No third-party transcription service without saying so explicitly, in the interface, at the moment of upload OPERATIONS - .env: DATABASE_PATH, STORAGE_PATH, BASE_URL, SESSION_SECRET, MODEL_PATH, RETENTION_DAYS - Migrations on boot, each once; nightly backup off the machine - Health endpoint reporting the transcription queue WHAT MATTERS MOST Consent, custom vocabulary and timestamp alignment. Build the upload path before the bot, and never let a summary stand without its source.
What you lose
- A meeting bot that joins the call, records it and posts the summary without anyone remembering to
- Speaker separation trained on real meetings with overlapping voices
- Live transcription during the call rather than after it
- Search across every meeting you have ever had
- Integrations that push notes into the calendar event and the team chat
If you would rather not build
- Fireflies — paid, meeting bot with CRM integrations
What it costs
read from their page 17 Aug 2026
| Plan | Billed monthly | Billed yearly | Last read |
|---|---|---|---|
| — | $16.99/mo | — | 17 Aug 2026 |
Their pricing page is where these came from. Seeing a different price? Tell us.
The escape hatch
open source · no votes, no paid placement
whisper.cpp
$0Runs speech-to-text locally on ordinary hardware, with word-level timestamps.
ggml-org/whisper.cppfree · open source
Vibe
$0Desktop app around Whisper with diarisation and subtitle export.
thewh1teagle/vibefree · open source
Why this verdict
our own opinion · changed only by a person
63/100
Verdict kinda at 63: local transcription is now genuinely good and the citation rule keeps the summary honest, which is more than most products do. The meeting bot is what makes Otter effortless, and building one is a consent problem before it is a technical one.
History
tracked since 9 Aug 2026 · nothing is ever overwritten
Questions about Otter.ai
answered from the record above
Is Otter.ai free?
No — the plan we track is $16.99 a month. Pro at $16.99/month billed monthly, around $8.33 annually, with a monthly transcription minute allowance.
Can you replace Otter.ai by building your own?
ALMOST. A weekend of work, and real gaps remain. Replacement score 63 out of 100, build time a weekend. Read what you lose before you decide.
How much does Otter.ai cost?
$16.99 a month on Pro — $203.88 a year. Recorded 9 Aug 2026.
What do you lose by replacing Otter.ai?
A meeting bot that joins the call, records it and posts the summary without anyone remembering to; Speaker separation trained on real meetings with overlapping voices; Live transcription during the call rather than after it; Search across every meeting you have ever had; Integrations that push notes into the calendar event and the team chat. If any of those carry weight for you, keep paying.
Is there an open-source alternative to Otter.ai?
Yes: whisper.cpp, Vibe. The prompt on this page is for when you want it your way instead.
Related entries
same category first, most replaced first
Every week, something stops being worth paying for.
New verdicts, prices that moved, entries added. One email a week. Unsubscribe in one click. Nothing is being sent yet — your address is kept here, and the first issue is the first thing it is used for.
free forever · no tracking pixel · stored here, never passed to anyone

