Tags: experimentation concept

Reading a Test Result

Date: 2026-08-16


A fixed order of checks, each of which can end the reading. The order isn’t preference — every step validates an assumption the next one depends on, so jumping to the primary metric means trusting four things you haven’t looked at.


What it is

Reading a test result is the sequence of validity checks performed before the outcome is believed, in an order where each step’s answer is a precondition for the next being meaningful.

The discipline is that any step can stop the reading entirely, and stopping is a complete answer.

The order

1  SRM                  failed → STOP. Nothing here is readable
2  Duration & sample     short → the p-value doesn't mean what it says
3  Guardrails         breached → the win is vetoed regardless
4  Primary metric              → read the interval, not the verdict
5  Everything else             → diagnostic only. Cannot change the decision

1. Sample Ratio Mismatch. The split matches the intended allocation, at p > 0.001. If it doesn’t, assignment or delivery is broken and the arms aren’t comparable — no analysis survives, and re-weighting doesn’t fix it. This is first because it’s cheap and it invalidates everything downstream.

2. Duration and sample. Did it run whole weeks, to the pre-registered endpoint, hitting the planned sample? A test stopped early or extended because of what it showed has an inflated false positive rate — Peeking, Test Duration.

3. Guardrail Metrics. Read these before the primary metric, deliberately. Reading them afterwards means reading them already knowing whether you want the test to pass, and the threshold gets negotiated.

4. The primary metric. One metric, the one in the pre-registration. Read the confidence interval and convert both bounds to money — see Practical vs Statistical Significance. The decision rule was written in advance; apply it.

5. Everything else. Secondary metrics explain why the primary moved. Segments generate hypotheses for next time. Neither can change this test’s verdict — The Multiple Comparisons Problem, Segmentation (test results).

In plain terms: the order exists so you find out the result is unreadable before you have an opinion about it. Once you’ve seen the primary metric, every subsequent check is being made by someone who wants a particular answer.

The four verdicts

Not two. Collapsing to win/lose is where most value is lost.

VerdictConditionAction
InvalidSRM, guardrail breach, or design brokenFix and rerun. Record it — these are real bugs affecting real users
ShipCI lower bound clears the pre-set thresholdShip, then validate
Don’t shipCI upper bound below the thresholdGenuine negative. Record the mechanism as disproven — Institutional Learning
InconclusiveInterval straddles the thresholdMore traffic, or a judgement call. Not a failure — Inconclusive Results

Before you write it up

Two diagnostics worth running on every result, both pre-registerable so they aren’t fishing:

  • New versus returning. Detects Novelty and Primacy Effects, which otherwise present as a real effect that later evaporates
  • Effect over time. A decaying or climbing trend means the same. Read the shape, not any single week

And one sanity check: does the mechanism in the hypothesis explain the pattern in the secondary metrics? If you predicted delivery cost was the barrier and add-to-cart moved but checkout entry didn’t, the result may be right and the explanation wrong — which matters more for the next five tests than this one’s verdict does.

Failure modes

  • Reading the primary metric first, then checking validity. Everything after is motivated
  • Treating “not significant” as “no effect” — Statistical Power, Inconclusive Results
  • Reading the p-value instead of the interval. The interval carries magnitude and precision; the p-value carries neither
  • Promoting a secondary metric because the primary didn’t win
  • Discovering a segment and reporting it as the finding
  • Not recording losses and inconclusives in the Experiment Archive. A programme that only records winners cannot compute its own win rate, and will overestimate itself permanently

Assembled as a procedure with the arithmetic in Guide - Statistics for CRO.