Tags: experimentation concept
Reading a Test Result
Date: 2026-08-16
A fixed order of checks, each of which can end the reading. The order isn’t preference — every step validates an assumption the next one depends on, so jumping to the primary metric means trusting four things you haven’t looked at.
What it is
Reading a test result is the sequence of validity checks performed before the outcome is believed, in an order where each step’s answer is a precondition for the next being meaningful.
The discipline is that any step can stop the reading entirely, and stopping is a complete answer.
The order
1 SRM failed → STOP. Nothing here is readable
2 Duration & sample short → the p-value doesn't mean what it says
3 Guardrails breached → the win is vetoed regardless
4 Primary metric → read the interval, not the verdict
5 Everything else → diagnostic only. Cannot change the decision
1. Sample Ratio Mismatch. The split matches the intended allocation, at p > 0.001. If it doesn’t, assignment or delivery is broken and the arms aren’t comparable — no analysis survives, and re-weighting doesn’t fix it. This is first because it’s cheap and it invalidates everything downstream.
2. Duration and sample. Did it run whole weeks, to the pre-registered endpoint, hitting the planned sample? A test stopped early or extended because of what it showed has an inflated false positive rate — Peeking, Test Duration.
3. Guardrail Metrics. Read these before the primary metric, deliberately. Reading them afterwards means reading them already knowing whether you want the test to pass, and the threshold gets negotiated.
4. The primary metric. One metric, the one in the pre-registration. Read the confidence interval and convert both bounds to money — see Practical vs Statistical Significance. The decision rule was written in advance; apply it.
5. Everything else. Secondary metrics explain why the primary moved. Segments generate hypotheses for next time. Neither can change this test’s verdict — The Multiple Comparisons Problem, Segmentation (test results).
In plain terms: the order exists so you find out the result is unreadable before you have an opinion about it. Once you’ve seen the primary metric, every subsequent check is being made by someone who wants a particular answer.
The four verdicts
Not two. Collapsing to win/lose is where most value is lost.
| Verdict | Condition | Action |
|---|---|---|
| Invalid | SRM, guardrail breach, or design broken | Fix and rerun. Record it — these are real bugs affecting real users |
| Ship | CI lower bound clears the pre-set threshold | Ship, then validate |
| Don’t ship | CI upper bound below the threshold | Genuine negative. Record the mechanism as disproven — Institutional Learning |
| Inconclusive | Interval straddles the threshold | More traffic, or a judgement call. Not a failure — Inconclusive Results |
Before you write it up
Two diagnostics worth running on every result, both pre-registerable so they aren’t fishing:
- New versus returning. Detects Novelty and Primacy Effects, which otherwise present as a real effect that later evaporates
- Effect over time. A decaying or climbing trend means the same. Read the shape, not any single week
And one sanity check: does the mechanism in the hypothesis explain the pattern in the secondary metrics? If you predicted delivery cost was the barrier and add-to-cart moved but checkout entry didn’t, the result may be right and the explanation wrong — which matters more for the next five tests than this one’s verdict does.
Failure modes
- Reading the primary metric first, then checking validity. Everything after is motivated
- Treating “not significant” as “no effect” — Statistical Power, Inconclusive Results
- Reading the p-value instead of the interval. The interval carries magnitude and precision; the p-value carries neither
- Promoting a secondary metric because the primary didn’t win
- Discovering a segment and reporting it as the finding
- Not recording losses and inconclusives in the Experiment Archive. A programme that only records winners cannot compute its own win rate, and will overestimate itself permanently
Assembled as a procedure with the arithmetic in Guide - Statistics for CRO.