Otter.ai

otter.aicontributed by Samuele Ongaro

ALMOST

A weekend of work, and real gaps remain.

Transcribes meetings and conversations: joins calls, separates speakers, produces a searchable transcript, and summarises the outcome.

Promptfree, for everyone, and the only version there is
Build me meeting transcription that replaces Otter — and read the two obstacles first.

**A bot that joins the call** so nobody has to remember to record is the feature, and building one means driving a meeting client programmatically, which each platform makes awkward in its own way. **Speaker separation with overlapping voices** is the other, and it is a research problem where local models are noticeably weaker. And before either: in many places recording a conversation requires everyone's consent, so the announcement is not a courtesy, it is the law.

STACK
- Node 20+ with Fastify
- SQLite through better-sqlite3, WAL mode, with FTS5
- whisper.cpp or an equivalent local model for transcription, and a diarisation model for speakers
- ffmpeg for audio handling
- Caddy in front

CONSENT, WHICH COMES FIRST
- Announce the recording at the start, audibly and visibly, before anything is captured
- Record that the announcement happened, with a timestamp, per meeting
- A way for anybody to object and stop it. Not a setting somebody has to find afterwards
- Write the jurisdictions you have checked in the README, and take advice. Recording a colleague without telling them is not a technical decision

THE DATA MODEL
- meetings: id, title, source, external_id, starts_at_utc, ends_at_utc, timezone, audio_path, duration_ms, status, consent_announced_at, created_by
- participants: id, meeting_id, name, email, speaker_label, joined_at, left_at
- segments: id, meeting_id, speaker_label, start_ms, end_ms, text, confidence, is_edited — the transcript, as segments rather than one blob, which is what makes search, editing and playback alignment possible
- summaries: id, meeting_id, kind, body_md, model_version, generated_at, is_edited
- actions: id, meeting_id, segment_id, text, assignee, due_at, status — extracted, then confirmed by a human
- highlights, comments, shares
- Audio kept as the original, with the transcript derived. Never discard the audio while the transcript is the only record, because transcription is imperfect and the recording is the evidence

GETTING THE AUDIO IN
- The easy path, and the one to build first: upload a recording, or point at a file the meeting platform already produced. That covers most of the value with none of the difficulty
- A browser recorder for an in-person conversation, capturing the microphone directly
- The bot that joins a call is the hard path: a headless client, an account, and per-platform behaviour that changes. Some platforms offer a proper API for this; check before writing anything, because if yours does, the problem disappears
- Live transcription during a call is a streaming model with partial results, which is a different pipeline from batch. Build batch first

TRANSCRIPTION
- Chunk the audio with overlap so a word split across a boundary is not lost, and stitch the results
- Store confidence per segment, and show low-confidence text differently. A transcript that hides its uncertainty gets quoted wrongly
- A custom vocabulary of names, products and jargon, applied as a post-pass — this is the single biggest quality improvement available and it costs almost nothing
- Speaker diarisation produces labels, not names. Let a human name each speaker once per meeting and remember the voice for next time, which is enough without solving the general problem

THE TRANSCRIPT
- Editable, with edits marked and the original kept
- Click a line to play from that moment. That alignment is what makes a transcript usable rather than just readable
- Search across every meeting with FTS5, with the matching line and the timestamp
- Export as text, Markdown, subtitles and a document

SUMMARIES, HONESTLY
- Generate from the transcript with a local model if you have one, and label it as generated
- Never present a summary without the transcript beside it. A summary is a lossy reading of a conversation people are accountable for, and the source must be one click away
- Action items extracted and then confirmed by a person before they become tasks. An automatically assigned action nobody agreed to is worse than none

PRIVACY
- This holds recordings of colleagues talking. Encrypt at rest, restrict access, log every playback, and set a real retention with a sweeper that deletes
- No third-party transcription service without saying so explicitly, in the interface, at the moment of upload

OPERATIONS
- .env: DATABASE_PATH, STORAGE_PATH, BASE_URL, SESSION_SECRET, MODEL_PATH, RETENTION_DAYS
- Migrations on boot, each once; nightly backup off the machine
- Health endpoint reporting the transcription queue

WHAT MATTERS MOST
Consent, custom vocabulary and timestamp alignment. Build the upload path before the bot, and never let a summary stand without its source.

What you lose

  • A meeting bot that joins the call, records it and posts the summary without anyone remembering to
  • Speaker separation trained on real meetings with overlapping voices
  • Live transcription during the call rather than after it
  • Search across every meeting you have ever had
  • Integrations that push notes into the calendar event and the team chat

If you would rather not build

  • Fireflies — paid, meeting bot with CRM integrations

What it costs

read from their page 17 Aug 2026

PlanBilled monthlyBilled yearlyLast read
—$16.99/mo—17 Aug 2026

Their pricing page is where these came from. Seeing a different price? Tell us.

The escape hatch

open source · no votes, no paid placement

whisper.cpp

$0

Runs speech-to-text locally on ordinary hardware, with word-level timestamps.

ggml-org/whisper.cppfree · open source

Vibe

$0

Desktop app around Whisper with diarisation and subtitle export.

thewh1teagle/vibefree · open source

Why this verdict

our own opinion · changed only by a person

63/100

Verdict kinda at 63: local transcription is now genuinely good and the citation rule keeps the summary honest, which is more than most products do. The meeting bot is what makes Otter effortless, and building one is a consent problem before it is a technical one.

History

tracked since 9 Aug 2026 · nothing is ever overwritten

Interest · last 30 dayspeak 2/day
views01230 Aug4 Sept9 Sept14 Sept19 Sept24 Sept28 Sept
— views— prompt copies none yet— votes none yet

Questions about Otter.ai

answered from the record above

Is Otter.ai free?

No — the plan we track is $16.99 a month. Pro at $16.99/month billed monthly, around $8.33 annually, with a monthly transcription minute allowance.

Can you replace Otter.ai by building your own?

ALMOST. A weekend of work, and real gaps remain. Replacement score 63 out of 100, build time a weekend. Read what you lose before you decide.

How much does Otter.ai cost?

$16.99 a month on Pro — $203.88 a year. Recorded 9 Aug 2026.

What do you lose by replacing Otter.ai?

A meeting bot that joins the call, records it and posts the summary without anyone remembering to; Speaker separation trained on real meetings with overlapping voices; Live transcription during the call rather than after it; Search across every meeting you have ever had; Integrations that push notes into the calendar event and the team chat. If any of those carry weight for you, keep paying.

Is there an open-source alternative to Otter.ai?

Yes: whisper.cpp, Vibe. The prompt on this page is for when you want it your way instead.

Related entries

same category first, most replaced first

All 40 in Notes, docs & writing

Not sending yet

Every week, something stops being worth paying for.

New verdicts, prices that moved, entries added. One email a week. Unsubscribe in one click. Nothing is being sent yet — your address is kept here, and the first issue is the first thing it is used for.

free forever · no tracking pixel · stored here, never passed to anyone

Esc