Tags: experimentation statistics concept

Stopping Rules

Date: 2026-08-17


The conditions that end a test, written down before it starts. Without them the test ends when someone looks at a dashboard and likes what they see — which is not a decision procedure, and it inflates the false positive rate to somewhere between three and five times the number you think you’re using.


A stopping rule is a condition, fixed before launch, that decides when a test ends and how its result is read.

Why “we’ll stop when it’s significant” can only end one way

Checking repeatedly and stopping at the first significant result means there is only one exit from the test. The false positive rate compounds with every look.

looks at the data      actual false positive rate (nominal α = 5%)

1                          5%
2                         ~8%
5                        ~14%
10                       ~19%
continuous monitoring    → approaches 100% given enough time

In plain terms: with a genuinely useless variant and enough patience, a peeking analyst will eventually see a “significant” result almost every time. The result is real in the sense that the arithmetic was done correctly; it just doesn’t mean what the 5% claims, because the 5% assumes one look at a pre-specified sample size — Peeking, P-Values.

The mechanism is that a test statistic wanders. Early on, with small samples, it wanders a long way. Stopping the moment it crosses a line selects for the wanders, not the effects.

What a stopping rule actually contains

Four conditions, all written before launch, all with numbers in them. Two abbreviations appear below: MDE, the minimum detectable effect — the smallest effect the test is designed to find — and SRM, sample ratio mismatch, where the observed split doesn’t match the intended one.

1  PLANNED END      sample size 210,000 per arm, or 3 full weeks —
                    whichever is LATER. computed at 3% baseline,
                    5% relative MDE, 80% power, 95% confidence

2  EARLY STOP FOR   only via the pre-declared sequential boundary
   SUCCESS          (or: not at all — this test runs to its end date)

3  EARLY STOP FOR   any guardrail breach beyond its threshold:
   HARM              · checkout completion −2% or worse
                     · page load p75 +200ms or worse
                     · error rate above 1%
                    → stop immediately, no discussion needed

4  ABANDON          if broken: SRM detected, tracking failure,
                    variant not rendering, or the change is reverted
                    for unrelated reasons

Condition 3 is not peeking, and this is the distinction that matters. Watching for harm and watching for success are asymmetric: you are not looking for a reason to declare a winner, you’re looking for a reason to protect users. The cost of a false alarm there is a cancelled test; the cost of missing real harm is real harm — Guardrail Metrics.

Condition 1’s “whichever is later” is doing real work. Hitting the sample in four days means you sampled four days of behaviour, not your customers — Test Duration.

Stopping early, legitimately

You can buy the right to look, but you must buy it before launch — it cannot be applied retrospectively to a test you’ve been watching.

ApproachWhat it costsWhen it fits
Fixed horizon, no looksNothing. The defaultMost tests. Two weeks is not long to wait
Group sequential (O’Brien-Fleming, Pocock)~5–15% more sample for the full runPre-planned interim analyses, e.g. at 33% and 66%
Always-valid inference / sequential~20–50% more sample to reach the same powerContinuous monitoring genuinely needed; dashboards anyone can read

The trade is always the same: the licence to look is paid for in power. A sequential design that lets you check daily will, on a test with no early winner, need meaningfully more traffic than a fixed-horizon design would have. That’s a fair price when peeking is unavoidable — and a waste when nobody was going to act early anyway — Sequential Testing, Always-Valid Inference.

The rules that make it hold

  • Write it in the test’s record before launch, with the numbers filled in — Pre-Registration
  • “Whichever is later”, never “whichever is sooner.” Sample and duration, both satisfied
  • Never extend a test because it’s nearly significant. Extending on the basis of the current result is peeking with extra steps, and it has exactly the same inflation. Extending for a pre-declared reason — a bank holiday fell inside the window — is fine, decided on the calendar rather than the p-value
  • Never restart a test that looked bad. “We fixed a small thing and reran it” resets the clock and discards the evidence, selectively
  • Nobody sees the primary metric mid-test unless the design permits it. Guardrails visible to everyone, primary metric hidden until the end, is a defensible and increasingly common setup
  • The person who wants the result should not be the person who decides it’s over. Separating those two roles removes most of the pressure that breaks stopping rules

What to do when it ends

The rule says when to stop, not what to conclude. Three outcomes, all valid:

interval entirely above 0     → ship, with the caveat that the
                                 true effect is probably smaller
                                 than measured — Winner's Curse

interval spans 0              → inconclusive. NOT "no effect".
                                 read the width — a tight interval
                                 around zero is informative, a wide
                                 one means you learned nothing
                                 Inconclusive Results

interval entirely below 0     → the change is harmful. this is a
                                 real finding and worth as much as a win

Two things to carry out of that table. A shipped winner’s true effect is almost always smaller than the one measured, because you selected it for crossing a line — Winner’s Curse. And the middle case is a distinct outcome with its own reading, not a failed version of the first — Inconclusive Results.

An inconclusive result is not a licence to run it longer. That’s the same violation, arriving after the fact.

Where it interacts

  • Peeking — the failure this note exists to prevent, and the arithmetic of it
  • Sample Size Calculation — supplies the number in condition 1, and a stopping rule without one is a preference
  • Test Duration — the calendar half of the same decision
  • Reading a Test Result — what happens after the rule fires
  • Multi-Armed Bandits — an alternative that reallocates continuously rather than stopping, and answers a different question