Tags: experimentation concept
Pre-Registration
Date: 2026-08-16
Writing the analysis down before the data exists. It’s what makes the p-value mean anything — because a p-value is a property of the procedure you committed to, not of the numbers you ended up with.
What it is
Pre-registration is recording the hypothesis, metrics, sample size, duration and analysis plan before a test launches, in a form that can’t be quietly edited afterwards.
It’s a commitment device, not documentation. The point is that the decisions were fixed at a moment when you couldn’t yet know which choice would flatter the result.
Why it’s load-bearing
The significance calculation assumes exactly one pre-specified comparison, evaluated once, at a pre-specified endpoint. Every degree of freedom you retain until after seeing data breaks that assumption:
| Freedom retained | Becomes |
|---|---|
| Which metric is primary | The Multiple Comparisons Problem |
| When to stop | Peeking |
| Which segment to report | Segmentation (test results) |
| Whether to exclude outliers, and above what | Choosing a threshold that produces the answer — Winsorisation and Capping |
| Whether to exclude a bad week | Same, with better rationalisation |
| One- or two-tailed | Halving the p-value after the fact — One-Tailed vs Two-Tailed Tests |
In plain terms: each of those choices is small, defensible in isolation, and made in good faith. Made after seeing the data, they collectively guarantee you can find a winner in noise. That’s The Garden of Forking Paths, and it doesn’t require anyone to be dishonest.
What goes in it
Short enough to actually write. A page:
HYPOTHESIS because we observed [X], we believe [change] will cause
[effect] for [audience], because [mechanism]
PRIMARY METRIC one. user-scoped conversion rate
GUARDRAILS revenue per visitor (−2%), LCP (+200ms), error rate
SECONDARY revenue per visitor, add-to-cart rate — diagnostic only
MDE 10% relative, derived from £15k build cost
(minimum detectable effect — the smallest the test can find)
SAMPLE 53,000 per arm
DURATION 9 weeks, Mon 1 Sep → Mon 3 Nov, whole weeks
UNIT user (cookie), 50/50
EXCLUSIONS bots, internal IPs, orders above £400 winsorised
(99th pct of last quarter)
SEGMENTS mobile / desktop, pre-registered. No others reported
STOPPING guardrail breach only. No early stop for success
DECISION RULE ship if CI lower bound exceeds +2% relative
The last line is the one most often missing and most valuable — deciding in advance what result would change what you do. It forces the Practical vs Statistical Significance conversation before anyone is emotionally invested.
Making it stick
A plan that can be edited after the fact isn’t a commitment. Options, cheapest first:
- Commit it to the repo with the test’s code. Timestamped by git, visible in the diff
- Post it in a channel where edits are visible, or send it to the stakeholders
- Record it in the experimentation tool if it locks fields on launch
- A dated document everyone has seen. Weakest, but far better than nothing
The bar isn’t cryptographic proof. It’s that changing it later requires an act someone would notice.
What it doesn’t prevent
Worth being clear so it isn’t over-claimed:
- It doesn’t stop you learning from surprises. A post-hoc finding is still a legitimate hypothesis — it just isn’t a result. Write it into the archive as the next test
- It doesn’t fix a bad design. Underpowered stays underpowered
- It doesn’t survive a change of question. If the test breaks and you rebuild it, that’s a new pre-registration
Failure modes
- Written after the build, so it rationalises decisions already made
- Vague primary metric — “conversion” without the denominator, timezone or window, so the specification doesn’t actually constrain anything. See Metric Design
- No decision rule, so “what do we do with a +1.5% non-significant result” gets argued in the readout
- Silently amended. The most damaging, because the artefact still looks like a commitment
- Treated as bureaucracy. It takes fifteen minutes and it’s the difference between a result and an anecdote