Tags: experimentation concept

Guardrail Metrics

Date: 2026-08-16


The things that must not get worse, whatever the primary metric does. They invert the usual logic — you’re looking for harm rather than gain, so a missed detection is the expensive error and the thresholds are set accordingly.


What it is

A guardrail metric is a metric monitored during a test that can stop or veto it, but can never make it a winner. It exists to catch damage the OEC wouldn’t show.

Three kinds, and they’re worth separating:

KindExampleCatches
Business harmRevenue per visitor, margin, returns rateA conversion win that costs money
Experience harmPage load, error rate, support contactsA change that breaks something
Trust harmUnsubscribes, complaints, review scoresA change that works now and costs later

The inverted logic

For the OEC you’re guarding against false positives — believing in an effect that isn’t there. For guardrails you’re guarding against false negatives: missing harm that is there.

That inversion changes everything about how they’re set:

  • Looser thresholds. A guardrail at 95% significance will miss real harm most of the time, because harm detection is exactly the underpowered case. Many teams run guardrails at 90% or use one-sided tests deliberately
  • Watched continuously. Checking guardrails daily is not Peeking, because you’re not looking at the outcome and you’re not claiming a discovery. Stopping early for harm is always legitimate
  • Non-inferiority framing. The useful question isn’t “did it get worse” but “can I rule out getting worse by more than X” — which is a confidence interval question, not a p-value question. See Non-Inferiority Tests

In plain terms: a guardrail that comes back “no significant harm” from an underpowered check has told you almost nothing. You need to know what size of harm it could have detected.

The interval is the answer

Worked. A test’s guardrail on revenue per visitor comes back non-significant, with a 95% interval of −6% to +4%.

annual revenue at risk    500,000 visitors × £1.50 = £750,000
interval lower bound      −6%  →  −£45,000/year

“No significant harm” here means the change could be costing £45,000 a year and this test couldn’t tell. If the primary metric’s win is worth £22,000, the guardrail hasn’t cleared it — it’s told you the downside risk exceeds the upside estimate.

That’s the guardrail doing its job, and it’s invisible if you only read the significance verdict. See Practical vs Statistical Significance.

Setting them

  • Nominate before launch, with a threshold. “Revenue per visitor must not drop more than 2%” is actionable; “watch revenue” is not
  • Keep the list short. Four or five. Every guardrail is another comparison, so a long list reintroduces The Multiple Comparisons Problem and produces stop-the-test noise
  • Include at least one you’d genuinely stop for. A guardrail nobody would act on is decoration
  • Same across the programme. Consistent guardrails make the archive comparable and mean the check becomes automatic rather than negotiated per test
  • Power them, or say you haven’t. A guardrail on a rare event — complaints, chargebacks — cannot detect anything at test scale. Record it as monitoring, not as a check that passed

The standing set

For a retail site, a defensible default:

STOP IMMEDIATELY          error rate, page unavailability, checkout failure
VETO A WIN                revenue per visitor, contribution margin, returns rate
FLAG FOR REVIEW           page load (LCP, INP), support contact rate,
                          add-to-cart-to-purchase ratio

Splitting them by what happens if breached is more useful than splitting by metric type, because it makes the decision automatic when the moment comes.

Failure modes

  • Guardrails checked only at the end, by which point the harm has run for a month
  • Treated as secondary metrics and used to claim a win. A guardrail can only veto
  • So many that breaches are routine, so everyone stops reacting
  • No threshold, so “worse” is argued about after the fact
  • Ignoring a breach because the primary won. This is the exact situation guardrails exist for, and the exact situation in which they’re most likely to be overruled