Feature flags with measurement attached: a flag rollout is treated as an experiment, with metrics, significance and automatic alerts if a treatment makes something worse.
Build me feature delivery with measurement, replacing Split — and know what makes this ALMOST. Flags are a weekend. The thing that earns the price is **rollouts that are measured rather than watched**: every release tied to metrics, with guardrails that halt it automatically when something degrades. That needs correct statistics, and correct statistics under continuous monitoring is where homemade versions produce confident wrong answers. STACK - Node 20+ with Fastify for the control plane - SQLite through better-sqlite3 for flags and definitions - Your existing warehouse or database for the metrics. Do not build an analytics store - A small SDK per language, evaluating locally - Caddy in front THE DATA MODEL - flags: id, key, kind, variations_json, default_variation, is_active - environments, and rules per flag per environment: position, condition_json, variation, rollout_json - rollouts: id, flag_id, environment_id, target_percent, started_at, current_percent, step_percent, step_interval_minutes, status, halted_at, halt_reason - assignments: id, flag_id, unit_id, variation, at — append-only, and the record every measurement depends on - metrics: id, name, kind, query_sql, unit_column, value_column, is_guardrail, direction, minimum_effect - results: id, rollout_id, metric_id, computed_at, variant_stats_json, method, verdict - Assignments must be recorded from the first minute. Everything else can be recomputed; that cannot EVALUATION - Deterministic bucketing: hash the unit key with the flag key and a salt. The same user always gets the same variation everywhere, with no coordination - The SDK holds the rule set in memory and evaluates locally. No network call in the request path - It loads from a local cache first, then updates. If the service is unreachable it keeps serving the last known rules indefinitely rather than falling back to defaults — that fallback is how a flag service takes your application down with it - Updates pushed over server-sent events with a slow poll as a fallback, so a kill switch takes effect in seconds - Percentage rollouts stable when the percentage changes: a unit in the first ten per cent stays in when it becomes twenty THE MEASURED ROLLOUT, WHICH IS THE PRODUCT - A rollout is a schedule: five per cent, then twenty, then fifty, then everybody, with an interval between steps - At each step, guardrail metrics are computed for the exposed group against the rest. If one degrades beyond its threshold, the rollout halts automatically and alerts. It does not roll back on its own unless you configure it to — but halting is automatic, always - Guardrails are the ones you always care about: errors, latency, crashes, conversion. Defined once and attached to every rollout by default - The primary metric is whatever the change is meant to improve, with a minimum effect stated in advance - A sample-ratio mismatch check at every step: if the exposure split is significantly off, the rollout is broken and the numbers mean nothing. This catches more real bugs than any other check THE STATISTICS - Use a published method valid under continuous monitoring — a sequential test, or a Bayesian approach with a stated prior. **Not** a fixed-horizon test looked at after every step, which will call a neutral change significant roughly a third of the time - Report an interval, never a point estimate, with the method named beside every number - Compute the sample size needed for the minimum effect before starting, and show progress towards it - Multiple metrics mean multiple comparisons; correct for it or say plainly that you have not DISCIPLINE - Every flag has an owner and an intended lifetime. Temporary flags are the point; four hundred permanent ones is a codebase nobody can reason about - A stale-flag report — fully rolled out, or unchanged for months — and a code search showing where each is referenced so removal is a small task - Every flag change audited with before and after values. A flag change is a production change OPERATIONS - .env: DATABASE_PATH, WAREHOUSE_URL, BASE_URL, SESSION_SECRET, SDK_KEY_SALT - Migrations on boot, each once; nightly backup off the machine - The control plane independent of the applications it serves, or they share an outage - Health endpoint reporting connected SDKs, rule version, and any rollout with a ratio mismatch WHAT MATTERS MOST Local evaluation that survives the service being down, and automatic halting on a guardrail. Turn the control plane off and confirm every application keeps serving; then degrade a metric on purpose and confirm the rollout stops.
What you lose
- Feature flags joined to metrics, so a rollout is measured rather than watched
- Statistical significance computed properly, including sequential testing
- Automatic alerting when a treatment degrades a guardrail metric
- SDKs for a dozen languages with local evaluation and no network in the hot path
- Audit trails a regulated team can show an auditor
If you would rather not build
- Flagsmith — open-source flags and remote config
What it costs
as published on their pricing page
| Plan | Billed monthly | Billed yearly | Last read |
|---|---|---|---|
| — | $33/mo | — | — |
Their pricing page is where these came from. Seeing a different price? Tell us.
The escape hatch
open source · no votes, no paid placement
Unleash
$0Feature flag server with gradual rollouts, segments and SDKs for most languages.
Unleash/unleashfree · open source
GrowthBook
$0Feature flags with experiment analysis run against your own warehouse.
growthbook/growthbookfree · open source
Why this verdict
our own opinion · changed only by a person
47/100
Verdict kinda at 47: flags with stable bucketing and local evaluation are a weekend. Measurement is where this gets its score capped — the statistics are easy to write and easy to get wrong, and the prompt spends its last section refusing to let you peek.
History
tracked since 9 Aug 2026 · nothing is ever overwritten
Questions about Split
answered from the record above
Is Split free?
No — the plan we track is $33 a month. Team plans start around $33 per seat per month billed annually; usage is metered on monthly tracked keys, and enterprise is quoted.
Can you replace Split by building your own?
ALMOST. A weekend of work, and real gaps remain. Replacement score 47 out of 100, build time a weekend. Read what you lose before you decide.
How much does Split cost?
$33 a month on Team — $396 a year. Recorded 9 Aug 2026.
What do you lose by replacing Split?
Feature flags joined to metrics, so a rollout is measured rather than watched; Statistical significance computed properly, including sequential testing; Automatic alerting when a treatment degrades a guardrail metric; SDKs for a dozen languages with local evaluation and no network in the hot path; Audit trails a regulated team can show an auditor. If any of those carry weight for you, keep paying.
Is there an open-source alternative to Split?
Yes: Unleash, GrowthBook. The prompt on this page is for when you want it your way instead.
Related entries
same category first, most replaced first
Every week, something stops being worth paying for.
New verdicts, prices that moved, entries added. One email a week. Unsubscribe in one click. Nothing is being sent yet — your address is kept here, and the first issue is the first thing it is used for.
free forever · no tracking pixel · stored here, never passed to anyone

