Tags: experimentation concept

Pre-Registration

Date: 2026-08-16


Writing the analysis down before the data exists. It’s what makes the p-value mean anything — because a p-value is a property of the procedure you committed to, not of the numbers you ended up with.


What it is

Pre-registration is recording the hypothesis, metrics, sample size, duration and analysis plan before a test launches, in a form that can’t be quietly edited afterwards.

It’s a commitment device, not documentation. The point is that the decisions were fixed at a moment when you couldn’t yet know which choice would flatter the result.

Why it’s load-bearing

The significance calculation assumes exactly one pre-specified comparison, evaluated once, at a pre-specified endpoint. Every degree of freedom you retain until after seeing data breaks that assumption:

Freedom retainedBecomes
Which metric is primaryThe Multiple Comparisons Problem
When to stopPeeking
Which segment to reportSegmentation (test results)
Whether to exclude outliers, and above whatChoosing a threshold that produces the answer — Winsorisation and Capping
Whether to exclude a bad weekSame, with better rationalisation
One- or two-tailedHalving the p-value after the fact — One-Tailed vs Two-Tailed Tests

In plain terms: each of those choices is small, defensible in isolation, and made in good faith. Made after seeing the data, they collectively guarantee you can find a winner in noise. That’s The Garden of Forking Paths, and it doesn’t require anyone to be dishonest.

What goes in it

Short enough to actually write. A page:

HYPOTHESIS      because we observed [X], we believe [change] will cause
                [effect] for [audience], because [mechanism]

PRIMARY METRIC  one. user-scoped conversion rate
GUARDRAILS      revenue per visitor (−2%), LCP (+200ms), error rate
SECONDARY       revenue per visitor, add-to-cart rate — diagnostic only

MDE             10% relative, derived from £15k build cost
                  (minimum detectable effect — the smallest the test can find)
SAMPLE          53,000 per arm
DURATION        9 weeks, Mon 1 Sep → Mon 3 Nov, whole weeks
UNIT            user (cookie), 50/50

EXCLUSIONS      bots, internal IPs, orders above £400 winsorised
                (99th pct of last quarter)
SEGMENTS        mobile / desktop, pre-registered. No others reported
STOPPING        guardrail breach only. No early stop for success

DECISION RULE   ship if CI lower bound exceeds +2% relative

The last line is the one most often missing and most valuable — deciding in advance what result would change what you do. It forces the Practical vs Statistical Significance conversation before anyone is emotionally invested.

Making it stick

A plan that can be edited after the fact isn’t a commitment. Options, cheapest first:

  • Commit it to the repo with the test’s code. Timestamped by git, visible in the diff
  • Post it in a channel where edits are visible, or send it to the stakeholders
  • Record it in the experimentation tool if it locks fields on launch
  • A dated document everyone has seen. Weakest, but far better than nothing

The bar isn’t cryptographic proof. It’s that changing it later requires an act someone would notice.

What it doesn’t prevent

Worth being clear so it isn’t over-claimed:

  • It doesn’t stop you learning from surprises. A post-hoc finding is still a legitimate hypothesis — it just isn’t a result. Write it into the archive as the next test
  • It doesn’t fix a bad design. Underpowered stays underpowered
  • It doesn’t survive a change of question. If the test breaks and you rebuild it, that’s a new pre-registration

Failure modes

  • Written after the build, so it rationalises decisions already made
  • Vague primary metric — “conversion” without the denominator, timezone or window, so the specification doesn’t actually constrain anything. See Metric Design
  • No decision rule, so “what do we do with a +1.5% non-significant result” gets argued in the readout
  • Silently amended. The most damaging, because the artefact still looks like a commitment
  • Treated as bureaucracy. It takes fifteen minutes and it’s the difference between a result and an anecdote