Feature flags and experimentation with the statistics done for you: rollouts, holdouts and results reported with confidence intervals rather than raw counts.
Build me experimentation that replaces Statsig — and know exactly where the difficulty is. **The statistics done correctly** is what you are paying for. Everything else — flags, assignment, a dashboard — is ordinary work. Getting the maths wrong does not produce an error; it produces a confident wrong answer that somebody ships. That asymmetry is why this is ALMOST rather than YES. STACK - Node 20+ with Fastify for the control plane - SQLite through better-sqlite3 for definitions, and your warehouse for the events - A small SDK per language, evaluating locally - Caddy in front THE DATA MODEL - gates and configs: the flag layer, with rules evaluated in order - experiments: id, key, hypothesis, variants_json, allocation_percent, unit_kind, layer_id, started_at, stopped_at, decision, decision_note - layers: id, name — mutually exclusive experiment slots, so two experiments cannot both change the same surface for the same user - assignments: id, experiment_id, unit_id, variant, at — append-only, and the record everything depends on - metrics: id, name, kind, definition_json, is_guardrail, direction, minimum_effect — count, sum, ratio, or a per-unit average - results: id, experiment_id, metric_id, computed_at, stats_json, method, verdict - Record assignments from the very first minute. Nothing else here can be reconstructed after the fact THE STATISTICS, WHICH IS THE JOB - Choose one published method and implement it faithfully. A sequential test valid under continuous monitoring, or a Bayesian approach with a stated prior. Never a fixed-horizon test that people look at daily — that combination calls a neutral change significant about a third of the time - Compute the required sample size before starting, from the baseline rate and the minimum effect worth acting on, and show progress towards it - Report an interval with the method named. A point estimate with a green tick is how a wrong decision gets made in a meeting - **Variance reduction** with a pre-experiment covariate is the single highest-value technique here: using each unit's own prior behaviour typically cuts the required sample size substantially. It is a published method, it is not difficult, and it turns a six-week experiment into a three-week one - Ratio metrics need the delta method, not a naive variance. Getting that wrong quietly understates uncertainty - Multiple metrics mean multiple comparisons. Correct for it, or state clearly that you have not and that the guardrails are not independent tests - Outliers: cap extreme per-unit values at a stated percentile before analysis, and say that you have. One customer's enormous order will otherwise decide your experiment CHECKS THAT CATCH REAL BUGS - **Sample ratio mismatch**: if the assignment split differs significantly from the intended one, the experiment is broken. This catches more genuine problems than any statistical refinement, and it should run on every experiment every day - Assignment before exposure: only count units that were actually exposed to the surface, not everybody bucketed. Otherwise the effect is diluted by people who never saw it - A pre-experiment period check: run the analysis on data from before the experiment started, where the answer should be no difference. If it is not, something in the pipeline is wrong ASSIGNMENT - Deterministic hashing of the unit key with the experiment key and a salt, so assignment is the same everywhere with no coordination - Layers so that overlapping experiments do not interfere, with the allocation split within a layer - The unit is a user, a device or an account — chosen per experiment and stated, because analysing a user-level experiment at the event level inflates significance dramatically DECISIONS - An experiment ends with a written decision and a reason, recorded. Without it the same question returns in six months - Keep stopped experiments and their assignments forever. That archive is where organisational learning lives, and it is the asset OPERATIONS - .env: DATABASE_PATH, WAREHOUSE_URL, BASE_URL, SESSION_SECRET - Migrations on boot, each once; nightly backup off the machine - Health endpoint reporting any experiment with a ratio mismatch WHAT MATTERS MOST A valid sequential method, the sample-ratio check, and variance reduction. Those three are the difference between an experimentation programme and an expensive way to confirm what somebody already believed.
What you lose
- Sequential testing and variance reduction done properly, which is where home-made experiments go wrong
- Automatic detection of a flag that is hurting a metric
- A single SDK covering flags, experiments and metrics
If you would rather not build
- A JSON file in your repository, for a handful of flags
What it costs
read from their page 15 Aug 2026
| Plan | Billed monthly | Billed yearly | Last read |
|---|---|---|---|
| — | $0/mo | — | 15 Aug 2026 |
Their pricing page is where these came from. Seeing a different price? Tell us.
The escape hatch
open source · no votes, no paid placement
Unleash
$0A self-hosted feature flag service with SDKs for most languages.
Unleash/unleashfree · open source
Flagsmith
$0Feature flags and remote config, self-hostable with an admin interface.
Flagsmith/flagsmithfree · open source
Why this verdict
our own opinion · changed only by a person
56/100
Verdict kinda at 56. Flags are easy and the prompt gives you a correct rollout. The statistics are where a home-made version quietly misleads you.
History
tracked since 10 Aug 2026 · nothing is ever overwritten
Questions about Statsig
answered from the record above
Is Statsig free?
No — the plan we track is $0 a month. Free up to a generous monthly event allowance; billing starts above that on events and sessions.
Can you replace Statsig by building your own?
ALMOST. A weekend of work, and real gaps remain. Replacement score 56 out of 100, build time a weekend. Read what you lose before you decide.
How much does Statsig cost?
$0 a month on Free tier — $0 a year. Recorded 10 Aug 2026.
What do you lose by replacing Statsig?
Sequential testing and variance reduction done properly, which is where home-made experiments go wrong; Automatic detection of a flag that is hurting a metric; A single SDK covering flags, experiments and metrics. If any of those carry weight for you, keep paying.
Is there an open-source alternative to Statsig?
Yes: Unleash, Flagsmith. The prompt on this page is for when you want it your way instead.
Related entries
same category first, most replaced first
Every week, something stops being worth paying for.
New verdicts, prices that moved, entries added. One email a week. Unsubscribe in one click. Nothing is being sent yet — your address is kept here, and the first issue is the first thing it is used for.
free forever · no tracking pixel · stored here, never passed to anyone

