Tags: experimentation web-dev concept

Rollouts as Experiments

Date: 2026-08-16


A staged rollout controls exposure. It does not, by itself, produce a comparison — and the sequential before/after reading everyone takes from one is confounded with time, traffic mix and everything else that changed that week.


What it is

A rollout becomes an experiment only when a comparable group is deliberately held back and measured concurrently. Ramping a feature from 1% to 100% is a release-safety practice; adding a held-out slice at each stage is what makes it measurable.

The distinction matters because the same tooling does both, so teams routinely believe they’ve run a test when they’ve run a release.

The three things people conflate

What it isFiled
Feature flagsThe mechanism: a runtime conditional that decouples deploy from releaseFeature Flags, Architecture
Progressive deliveryThe release practice: canary, percentage ramp, ring deployment, automatic rollbackProgressive Delivery, Architecture
Rollouts as experimentsThe measurement question sitting on top of both: when does a ramp support a causal claimHere

The first two answer how do I ship this safely. This one answers can I learn anything from having shipped it that way.

Why the naive reading fails

The tempting analysis is before/after: conversion was 3.00% last week at 0% rollout, it’s 3.12% this week at 50%. That comparison is worthless, and for the same reason every before/after is:

SEQUENTIAL (what a ramp gives you by default)

  week 1   0% exposed    3.00%   ← different week, different traffic,
  week 2  25% exposed    3.06%     different campaigns, different weather
  week 3  50% exposed    3.12%
  week 4 100% exposed    3.09%

CONCURRENT (what makes it an experiment)

  week 3   50% exposed   3.12%   ← same week, same traffic, same everything
           50% held back 3.00%     except the feature

The sequential reading is confounded with time. Campaigns changed, the mix shifted, a competitor ran a sale. It also can’t survive Regression to the Mean if the rollout was triggered by a bad number, and it has no way to detect Sample Ratio Mismatch because there’s no ratio to check.

In plain terms: comparing this week to last week compares two different worlds. Comparing two groups in the same week compares one world to itself.

Making a rollout measurable

  • Hold back a slice at every stage, not just early. A 50% ramp is a 50/50 test for free; a 95% ramp with 5% held back is still a valid comparison, just an underpowered one. Ramping to 100% ends the experiment
  • Randomise the holdback, don’t select it. “Everyone except internal users” or “all regions except one” is not random, and both alternatives differ systematically — see Why Randomisation Works
  • Log exposure, not just assignment. A user in the treated group who never reached the feature dilutes the effect. See Experiment Assignment Tracking
  • Keep the ratio stable long enough to power it. Ramping 1% → 5% → 25% weekly means each stage has its own sample size, and the early stages have almost none
  • Don’t pool across stages. Combining the 1% week with the 50% week compares different populations at different times, which reintroduces the confounding you were avoiding

When a ramp is genuinely not an experiment

Being honest about this is more useful than pretending otherwise. A ramp is a release practice, not a test, when:

  • The feature must ship regardless. A compliance change, a security fix, a supplier migration. There’s no decision the measurement would inform
  • The holdback is unethical or impractical — you can’t withhold a bug fix
  • The ramp is too fast to power anything. Hours at each stage measures nothing
  • You’re watching for breakage, not effect. That’s monitoring, and it’s a legitimate and different job — see Guardrail Metrics

In those cases say “we rolled it out safely” rather than “it improved conversion 4%”. The second claim is the one that ends up in a deck and then in someone’s belief system.

The long-run version

For changes too small to test individually, the same idea scaled up: keep a permanent randomly-assigned slice of users excluded from everything shipped this quarter, and compare at the end. That measures the cumulative effect of the whole programme, which is often the only way to detect the aggregate of many sub-MDE changes — see Holdout Groups and Minimum Detectable Effect.