Tags: experimentation concept

Secondary and Diagnostic Metrics

Date: 2026-08-17


Everything measured that isn’t the primary metric and isn’t a guardrail. Their job is explaining why a result happened; the discipline is that they never get to decide whether it happened, because a metric that can overturn the verdict has become a second primary one.


The three roles

PrimaryGuardrailSecondary / diagnostic
QuestionDid it work?Did it break anything?Why did that happen?
CountExactly oneA fixed standing setAs many as useful
DecidedBefore launchBefore launch, once, for all testsCan be added freely
PoweredYes, deliberatelyEnough to catch large harmUsually not, and that’s fine
Can change the ship decisionYesYes — it can vetoNo
Multiplicity correctionNot needed (n=1)Applied to the setNot applied — see below

The distinction between guardrail and secondary is the one that gets blurred. A guardrail can stop a ship; a diagnostic cannot. If you find yourself arguing that a diagnostic should block the rollout, it was a guardrail and should have been named as one before launch — Guardrail Metrics.

What they’re for

Reading a result is not just “did the number move” — it’s “does the story hold together”. Diagnostics are how you check that.

test: added delivery-date estimates to the product page
primary: checkout completion rate      +4.1%   (95% CI +1.2% to +7.1%)

diagnostics
  product page → add to cart           +5.8%   ← the change is acting where expected
  add to cart → checkout start         +0.3%   ← unchanged, as predicted
  checkout start → completion          −0.4%   ← unchanged
  delivery-estimate element viewed     91%     ← the change was actually seen
  returns rate (7d, early)             −2.0%   ← consistent with better expectations

story: people who saw a delivery date added to basket more.
       nothing downstream changed. the mechanism is the one we hypothesised.

Now the same primary result with diagnostics that don’t cohere:

primary: checkout completion rate      +4.1%
  product page → add to cart           −0.2%   ← the change did nothing here
  delivery-estimate element viewed     31%     ← two-thirds never saw it
  sessions per user                    −6.0%   ← fewer visits, same buyers

story: the primary moved without the mechanism moving. something else
       is going on — a tracking change, a traffic shift, or noise.
       do not ship on this.

Same headline number, opposite conclusion. The diagnostics didn’t overturn the primary metric — they revealed that the primary metric probably isn’t measuring what the hypothesis claimed, which is a different and legitimate objection.

The multiplicity question, answered properly

Twenty diagnostics at 95% confidence means roughly one false positive per test, by construction:

P(at least one false positive) = 1 − 0.95²⁰ = 1 − 0.358 = 0.642

In plain terms: with twenty diagnostic metrics, there’s about a 64% chance at least one shows a “significant” movement that isn’t real. Nearly two tests in three will hand you an exciting spurious finding — The Multiple Comparisons Problem.

The resolution isn’t to correct them — it’s to stop treating them as decisions. A correction is needed when you’re making a claim; diagnostics aren’t claims, they’re context. So:

  • No correction applied, because no decision depends on any single one
  • Read them as a pattern, not individually. Does the set tell a coherent story consistent with the hypothesis?
  • Never report a diagnostic as a finding. “The test won, and it also improved mobile newsletter signups by 9%” is exactly the sentence that turns noise into a believed fact
  • A surprising diagnostic is a hypothesis for the next test, never a result from this one — Hypothesis Design

Choosing them

Pick them from the hypothesis, before launch, by asking “if my proposed mechanism is real, what else must be true?”

  • Funnel steps around the change — where the effect should appear, and where it shouldn’t
  • Exposure and engagement with the changed element — did anyone actually see it? A null result with 30% viewing is a delivery problem, not a design one
  • The step immediately downstream — to catch effects that shift behaviour rather than adding it
  • Counts as well as rates. A rate can improve because the denominator fell, which is a completely different event — Metric Design
  • Anything that would falsify the mechanism. The most valuable diagnostic is the one you’d be uncomfortable to see move

Add them at analysis time freely. Unlike the primary metric, there’s no integrity cost to computing a new diagnostic after the fact — precisely because it can’t change the verdict. This is the practical benefit of keeping the roles strictly separate.

Failure modes

  • Promotion after the fact. The primary is flat, a secondary is up, the secondary becomes the story. This is the main way tests get “won” and it is P-Hacking with better manners
  • The dashboard with forty metrics and no marked primary. Someone will find a green one
  • Diagnostics read at the same confidence as the primary, giving them unearned authority. Report them as directional, with intervals, and say plainly that they’re underpowered
  • Segmented diagnostics, which multiply the comparisons again and are the most reliable source of false discoveries in the whole practice — Segmentation (test results)
  • No diagnostics at all, which leaves you unable to distinguish a real win from a measurement artefact, and unable to learn anything transferable from a null — Institutional Learning

Where it interacts