Tags: statistics experimentation concept

Peeking

Date: 2026-08-16


Checking a fixed-horizon test in flight and stopping when it looks good. Every look is another chance to cross the line by luck, so the false positive rate you designed for is not the one you’re running at — and the report still prints 5%.


What it is

Peeking is evaluating a test’s primary metric before it reaches its planned sample size, and letting what you see influence when you stop.

It isn’t the looking that breaks it — it’s that looking and stopping together form a different decision rule than the one the test was designed around.

The mechanism

A fixed-horizon test is built on one assumption: you evaluate it once, at the planned sample size. The 5% false positive rate is calculated for exactly one comparison against the threshold.

A test’s observed difference wanders as data accumulates. Early on the sample is small, the standard error is large, and the wandering is wide. Given enough looks, that random walk will cross the significance line at some point in almost any test — including one where nothing is happening.

day  3   +8.2%   p = 0.04   ← "significant!" stop here and you ship noise
day  7   +2.1%   p = 0.31
day 14   −0.4%   p = 0.88
day 21   +1.2%   p = 0.44
day 28   +0.6%   p = 0.71   ← the planned endpoint. Nothing here.

Nothing changed on the site across those four weeks. The day-3 reading is not an early sighting of an effect; it’s the widest part of the random walk.

How much it costs

Five looks, treating each naively as an independent chance at 5%:

P(no false positive) = 0.95⁵    = 0.774
P(at least one)      = 1 − 0.774 = 22.6%

Looks at a running test are correlated rather than independent — day 14’s data contains day 7’s — so the true inflation is lower than 22.6%. But the direction and rough scale hold: you designed a 5% error rate and you’re operating somewhere around three times that. Continuous daily monitoring over a fortnight is worse again. [CHECK: exact inflation figures for common monitoring patterns against a simulation source before quoting precisely.]

In plain terms: stopping the moment a test looks significant means you have selected for the moment it looked best. That moment arrives in tests where nothing is happening, and it arrives earliest in them, because small samples swing furthest.

Why the intuition fails

The instinct is that looking doesn’t change anything — the data is the data, and observing it is passive.

That’s true of the data and false of the decision rule. “Run to 106,000 visitors then evaluate” and “evaluate daily and stop at the first significant result” are different procedures with different error rates, applied to the same data. The p-value is a property of the procedure, not of the numbers. Change the procedure, and the printed p-value stops describing it.

The second failure: an early significant result feels like strong evidence because the effect looks large. It looks large because the sample is small — see Winner’s Curse and Effect Size on why early winners are inflated.

What to do instead

  • Commit to the endpoint in writing before launch, and to a minimum duration of at least one full week regardless of sample — Test Duration, Pre-Registration
  • If you must monitor continuously, use a design built for it. Sequential Testing and Always-Valid Inference spend some power to buy the licence to look. You cannot retrofit that licence onto a fixed-horizon test
  • Separate monitoring from evaluation. Check daily for Sample Ratio Mismatch, broken variants and guardrail harm. Do not look at the primary metric
  • Hide the primary metric in the dashboard if the temptation is organisational rather than personal. This is a normal and effective control

The legitimate exception

Stopping early for harm is fine and correct. The asymmetry is deliberate: you are not trying to claim a discovery, you’re avoiding damage, and a false alarm costs you a test slot rather than a wrong belief. Set guardrail stopping rules in advance, watch them continuously, and stop without hesitation — Guardrail Metrics, Stopping Rules.