Tags: experimentation web-dev concept
Rollouts as Experiments
Date: 2026-08-16
A staged rollout controls exposure. It does not, by itself, produce a comparison — and the sequential before/after reading everyone takes from one is confounded with time, traffic mix and everything else that changed that week.
What it is
A rollout becomes an experiment only when a comparable group is deliberately held back and measured concurrently. Ramping a feature from 1% to 100% is a release-safety practice; adding a held-out slice at each stage is what makes it measurable.
The distinction matters because the same tooling does both, so teams routinely believe they’ve run a test when they’ve run a release.
The three things people conflate
| What it is | Filed | |
|---|---|---|
| Feature flags | The mechanism: a runtime conditional that decouples deploy from release | Feature Flags, Architecture |
| Progressive delivery | The release practice: canary, percentage ramp, ring deployment, automatic rollback | Progressive Delivery, Architecture |
| Rollouts as experiments | The measurement question sitting on top of both: when does a ramp support a causal claim | Here |
The first two answer how do I ship this safely. This one answers can I learn anything from having shipped it that way.
Why the naive reading fails
The tempting analysis is before/after: conversion was 3.00% last week at 0% rollout, it’s 3.12% this week at 50%. That comparison is worthless, and for the same reason every before/after is:
SEQUENTIAL (what a ramp gives you by default)
week 1 0% exposed 3.00% ← different week, different traffic,
week 2 25% exposed 3.06% different campaigns, different weather
week 3 50% exposed 3.12%
week 4 100% exposed 3.09%
CONCURRENT (what makes it an experiment)
week 3 50% exposed 3.12% ← same week, same traffic, same everything
50% held back 3.00% except the feature
The sequential reading is confounded with time. Campaigns changed, the mix shifted, a competitor ran a sale. It also can’t survive Regression to the Mean if the rollout was triggered by a bad number, and it has no way to detect Sample Ratio Mismatch because there’s no ratio to check.
In plain terms: comparing this week to last week compares two different worlds. Comparing two groups in the same week compares one world to itself.
Making a rollout measurable
- Hold back a slice at every stage, not just early. A 50% ramp is a 50/50 test for free; a 95% ramp with 5% held back is still a valid comparison, just an underpowered one. Ramping to 100% ends the experiment
- Randomise the holdback, don’t select it. “Everyone except internal users” or “all regions except one” is not random, and both alternatives differ systematically — see Why Randomisation Works
- Log exposure, not just assignment. A user in the treated group who never reached the feature dilutes the effect. See Experiment Assignment Tracking
- Keep the ratio stable long enough to power it. Ramping 1% → 5% → 25% weekly means each stage has its own sample size, and the early stages have almost none
- Don’t pool across stages. Combining the 1% week with the 50% week compares different populations at different times, which reintroduces the confounding you were avoiding
When a ramp is genuinely not an experiment
Being honest about this is more useful than pretending otherwise. A ramp is a release practice, not a test, when:
- The feature must ship regardless. A compliance change, a security fix, a supplier migration. There’s no decision the measurement would inform
- The holdback is unethical or impractical — you can’t withhold a bug fix
- The ramp is too fast to power anything. Hours at each stage measures nothing
- You’re watching for breakage, not effect. That’s monitoring, and it’s a legitimate and different job — see Guardrail Metrics
In those cases say “we rolled it out safely” rather than “it improved conversion 4%”. The second claim is the one that ends up in a deck and then in someone’s belief system.
The long-run version
For changes too small to test individually, the same idea scaled up: keep a permanent randomly-assigned slice of users excluded from everything shipped this quarter, and compare at the end. That measures the cumulative effect of the whole programme, which is often the only way to detect the aggregate of many sub-MDE changes — see Holdout Groups and Minimum Detectable Effect.