Tags: experimentation statistics concept
Metric Selection for Tests
Date: 2026-08-17
Choosing what a test measures, which is a trade between how much you care about a metric and how much traffic it needs. The metric the business cares most about — revenue per visitor — is usually the one it cannot afford to test on, and pretending otherwise produces a year of inconclusive tests.
Metric selection is choosing, before launch, the primary metric a test’s decision rests on, plus the secondary and guardrail metrics reported alongside it.
The trade
meaningfulness →
low high
┌─────────────────────┬─────────────────────┐
high │ clicks on the │ ★ add-to-cart rate │
│ hero button │ checkout starts │
sensi- │ │ step completion │
tivity │ "engagement" │ │
├─────────────────────┼─────────────────────┤
│ scroll depth │ revenue per visitor │
low │ time on page │ lifetime value │
│ │ retention at 90d │
└─────────────────────┴─────────────────────┘
↑
what everyone asks for, and
what almost nobody can power
The top-right cell is where testable metrics live. It’s a small box, and finding what belongs in it for your site is most of the work.
Why revenue per visitor is so expensive
Sensitivity is driven by variance relative to the mean — the coefficient of variation, the standard deviation divided by the mean. Revenue per visitor has an enormous one, because most visitors spend nothing and a few spend a great deal.
CONVERSION RATE REVENUE PER VISITOR
values: 0 or 1 values: 0, 0, 0, 0, £34, 0, 0, £890, 0…
mean: 0.03 mean: £2.40
sd: √(0.03 × 0.97) = 0.171 sd: ~£28 (one order in a long tail)
CV: 0.171 / 0.03 = 5.7 CV: 28 / 2.40 = 11.7
~2× the CV → ~4× the sample,
because sample scales with CV²
In plain terms: to detect the same relative change, revenue per visitor typically needs somewhere around four times the traffic of conversion rate — sometimes far more, because a single large order can move the mean of an arm on its own. That’s the whole reason the metric everyone wants is the metric nobody powers — Metric Sensitivity, Skewed and Heavy-Tailed Distributions.
Worked, with real numbers. 3% baseline conversion, detecting a 5% relative lift at 80% power and 95% confidence:
conversion rate ≈ 210,000 visitors per arm
revenue per visitor ≈ 840,000 per arm at 4× — more if the tail is heavy
at 40,000 visitors/week: conversion → 10.5 weeks
revenue/visitor → 42 weeks
Forty-two weeks is not a test. It’s a year in which the site, the traffic mix and the product have all changed — Test Duration.
Making the expensive metric affordable
Before giving up on it, three levers that genuinely work:
- Winsorisation and Capping. Cap order values at, say, the 99th percentile. One £8,000 trade order shouldn’t decide a homepage test. This routinely cuts the variance substantially and biases the estimate only in the tail you capped
- Variance Reduction, usually CUPED — controlled experiment using pre-experiment data. Subtract each user’s pre-test spend from their outcome. With a pre-period correlation of ρ = 0.5 the variance falls by ρ² = 25%, which is a 25% reduction in required sample for free
- A ratio with a better denominator. Revenue per purchaser has far lower variance than revenue per visitor — but it’s conditioned on an outcome the test may have changed, so it can’t be a primary metric on its own — see the trap below
The surrogate trap
The standard move when the real metric is unaffordable is to test a closer, cheaper one — add-to-cart instead of purchase. It’s usually right, and it has a specific failure mode.
A surrogate is only valid if the treatment can’t move the surrogate without moving the outcome in the same direction. That condition fails constantly:
change: a large, prominent "Add to basket" button on the listing page
add-to-cart rate +18% ← the surrogate says ship it
checkout starts +2%
orders −1% ← people added the wrong size, then abandoned
revenue per visitor −3%
The surrogate moved because the change made adding easier, not because it made buying more likely. This is Goodhart’s law arriving inside a two-week test.
How to use a surrogate safely:
- Validate it historically. Across your archive of past tests, does the surrogate’s movement predict the outcome’s? If you have twenty tests with both measured, that’s an answerable question — and answering it is one of the best uses of an Experiment Archive
- Keep the real metric as a guardrail, underpowered but watched. You can’t detect a 2% loss, but you can detect a 15% collapse — Guardrail Metrics
- Prefer a surrogate downstream of the change, not upstream. Checkout starts is a better surrogate than button clicks because more of the funnel has to survive
The rules
- One primary metric, named before launch — Overall Evaluation Criterion
- Pick it from the sensitivity/meaningfulness box, not from the org chart. “The CFO wants revenue” is not a power calculation
- Ratio metrics need the right denominator and the right variance formula. Revenue per visitor is a ratio of two random quantities, and its standard error is not the naive one — Ratio Metrics
- Never change the primary metric after launch. If the chosen one turns out to be unpowered, the test is inconclusive and the next one is designed better — Pre-Registration
- Write down what result would make you ship, in the metric’s units, before you see any data. If no plausible result changes the decision, the metric is wrong or the test is unnecessary
Where it interacts
- Metric Sensitivity — the statistics underneath this, and the ranking of common metrics by cost
- Metric Design — the definition itself, which has to be settled before it can be tested on
- Secondary and Diagnostic Metrics — everything that explains the result without being allowed to decide it
- Practical vs Statistical Significance — a detectable effect on a cheap metric can still be worth nothing in pounds