Tags: experimentation concept
Secondary and Diagnostic Metrics
Date: 2026-08-17
Everything measured that isn’t the primary metric and isn’t a guardrail. Their job is explaining why a result happened; the discipline is that they never get to decide whether it happened, because a metric that can overturn the verdict has become a second primary one.
The three roles
| Primary | Guardrail | Secondary / diagnostic | |
|---|---|---|---|
| Question | Did it work? | Did it break anything? | Why did that happen? |
| Count | Exactly one | A fixed standing set | As many as useful |
| Decided | Before launch | Before launch, once, for all tests | Can be added freely |
| Powered | Yes, deliberately | Enough to catch large harm | Usually not, and that’s fine |
| Can change the ship decision | Yes | Yes — it can veto | No |
| Multiplicity correction | Not needed (n=1) | Applied to the set | Not applied — see below |
The distinction between guardrail and secondary is the one that gets blurred. A guardrail can stop a ship; a diagnostic cannot. If you find yourself arguing that a diagnostic should block the rollout, it was a guardrail and should have been named as one before launch — Guardrail Metrics.
What they’re for
Reading a result is not just “did the number move” — it’s “does the story hold together”. Diagnostics are how you check that.
test: added delivery-date estimates to the product page
primary: checkout completion rate +4.1% (95% CI +1.2% to +7.1%)
diagnostics
product page → add to cart +5.8% ← the change is acting where expected
add to cart → checkout start +0.3% ← unchanged, as predicted
checkout start → completion −0.4% ← unchanged
delivery-estimate element viewed 91% ← the change was actually seen
returns rate (7d, early) −2.0% ← consistent with better expectations
story: people who saw a delivery date added to basket more.
nothing downstream changed. the mechanism is the one we hypothesised.
Now the same primary result with diagnostics that don’t cohere:
primary: checkout completion rate +4.1%
product page → add to cart −0.2% ← the change did nothing here
delivery-estimate element viewed 31% ← two-thirds never saw it
sessions per user −6.0% ← fewer visits, same buyers
story: the primary moved without the mechanism moving. something else
is going on — a tracking change, a traffic shift, or noise.
do not ship on this.
Same headline number, opposite conclusion. The diagnostics didn’t overturn the primary metric — they revealed that the primary metric probably isn’t measuring what the hypothesis claimed, which is a different and legitimate objection.
The multiplicity question, answered properly
Twenty diagnostics at 95% confidence means roughly one false positive per test, by construction:
P(at least one false positive) = 1 − 0.95²⁰ = 1 − 0.358 = 0.642
In plain terms: with twenty diagnostic metrics, there’s about a 64% chance at least one shows a “significant” movement that isn’t real. Nearly two tests in three will hand you an exciting spurious finding — The Multiple Comparisons Problem.
The resolution isn’t to correct them — it’s to stop treating them as decisions. A correction is needed when you’re making a claim; diagnostics aren’t claims, they’re context. So:
- No correction applied, because no decision depends on any single one
- Read them as a pattern, not individually. Does the set tell a coherent story consistent with the hypothesis?
- Never report a diagnostic as a finding. “The test won, and it also improved mobile newsletter signups by 9%” is exactly the sentence that turns noise into a believed fact
- A surprising diagnostic is a hypothesis for the next test, never a result from this one — Hypothesis Design
Choosing them
Pick them from the hypothesis, before launch, by asking “if my proposed mechanism is real, what else must be true?”
- Funnel steps around the change — where the effect should appear, and where it shouldn’t
- Exposure and engagement with the changed element — did anyone actually see it? A null result with 30% viewing is a delivery problem, not a design one
- The step immediately downstream — to catch effects that shift behaviour rather than adding it
- Counts as well as rates. A rate can improve because the denominator fell, which is a completely different event — Metric Design
- Anything that would falsify the mechanism. The most valuable diagnostic is the one you’d be uncomfortable to see move
Add them at analysis time freely. Unlike the primary metric, there’s no integrity cost to computing a new diagnostic after the fact — precisely because it can’t change the verdict. This is the practical benefit of keeping the roles strictly separate.
Failure modes
- Promotion after the fact. The primary is flat, a secondary is up, the secondary becomes the story. This is the main way tests get “won” and it is P-Hacking with better manners
- The dashboard with forty metrics and no marked primary. Someone will find a green one
- Diagnostics read at the same confidence as the primary, giving them unearned authority. Report them as directional, with intervals, and say plainly that they’re underpowered
- Segmented diagnostics, which multiply the comparisons again and are the most reliable source of false discoveries in the whole practice — Segmentation (test results)
- No diagnostics at all, which leaves you unable to distinguish a real win from a measurement artefact, and unable to learn anything transferable from a null — Institutional Learning
Where it interacts
- Overall Evaluation Criterion — the one metric these exist to explain rather than compete with
- Reading a Test Result — where in the ordered sequence diagnostics get read, which is last
- Inconclusive Results — diagnostics are what make a null test informative instead of a shrug
- Experiment Archive — diagnostics are the part worth keeping, because they’re what makes a past test transferable to a new question