Tags: experimentation concept
Guardrail Metrics
Date: 2026-08-16
The things that must not get worse, whatever the primary metric does. They invert the usual logic — you’re looking for harm rather than gain, so a missed detection is the expensive error and the thresholds are set accordingly.
What it is
A guardrail metric is a metric monitored during a test that can stop or veto it, but can never make it a winner. It exists to catch damage the OEC wouldn’t show.
Three kinds, and they’re worth separating:
| Kind | Example | Catches |
|---|---|---|
| Business harm | Revenue per visitor, margin, returns rate | A conversion win that costs money |
| Experience harm | Page load, error rate, support contacts | A change that breaks something |
| Trust harm | Unsubscribes, complaints, review scores | A change that works now and costs later |
The inverted logic
For the OEC you’re guarding against false positives — believing in an effect that isn’t there. For guardrails you’re guarding against false negatives: missing harm that is there.
That inversion changes everything about how they’re set:
- Looser thresholds. A guardrail at 95% significance will miss real harm most of the time, because harm detection is exactly the underpowered case. Many teams run guardrails at 90% or use one-sided tests deliberately
- Watched continuously. Checking guardrails daily is not Peeking, because you’re not looking at the outcome and you’re not claiming a discovery. Stopping early for harm is always legitimate
- Non-inferiority framing. The useful question isn’t “did it get worse” but “can I rule out getting worse by more than X” — which is a confidence interval question, not a p-value question. See Non-Inferiority Tests
In plain terms: a guardrail that comes back “no significant harm” from an underpowered check has told you almost nothing. You need to know what size of harm it could have detected.
The interval is the answer
Worked. A test’s guardrail on revenue per visitor comes back non-significant, with a 95% interval of −6% to +4%.
annual revenue at risk 500,000 visitors × £1.50 = £750,000
interval lower bound −6% → −£45,000/year
“No significant harm” here means the change could be costing £45,000 a year and this test couldn’t tell. If the primary metric’s win is worth £22,000, the guardrail hasn’t cleared it — it’s told you the downside risk exceeds the upside estimate.
That’s the guardrail doing its job, and it’s invisible if you only read the significance verdict. See Practical vs Statistical Significance.
Setting them
- Nominate before launch, with a threshold. “Revenue per visitor must not drop more than 2%” is actionable; “watch revenue” is not
- Keep the list short. Four or five. Every guardrail is another comparison, so a long list reintroduces The Multiple Comparisons Problem and produces stop-the-test noise
- Include at least one you’d genuinely stop for. A guardrail nobody would act on is decoration
- Same across the programme. Consistent guardrails make the archive comparable and mean the check becomes automatic rather than negotiated per test
- Power them, or say you haven’t. A guardrail on a rare event — complaints, chargebacks — cannot detect anything at test scale. Record it as monitoring, not as a check that passed
The standing set
For a retail site, a defensible default:
STOP IMMEDIATELY error rate, page unavailability, checkout failure
VETO A WIN revenue per visitor, contribution margin, returns rate
FLAG FOR REVIEW page load (LCP, INP), support contact rate,
add-to-cart-to-purchase ratio
Splitting them by what happens if breached is more useful than splitting by metric type, because it makes the decision automatic when the moment comes.
Failure modes
- Guardrails checked only at the end, by which point the harm has run for a month
- Treated as secondary metrics and used to claim a win. A guardrail can only veto
- So many that breaches are routine, so everyone stops reacting
- No threshold, so “worse” is argued about after the fact
- Ignoring a breach because the primary won. This is the exact situation guardrails exist for, and the exact situation in which they’re most likely to be overruled