Tags: experimentation statistics concept

Metric Selection for Tests

Date: 2026-08-17


Choosing what a test measures, which is a trade between how much you care about a metric and how much traffic it needs. The metric the business cares most about — revenue per visitor — is usually the one it cannot afford to test on, and pretending otherwise produces a year of inconclusive tests.


Metric selection is choosing, before launch, the primary metric a test’s decision rests on, plus the secondary and guardrail metrics reported alongside it.

The trade

                    meaningfulness  →

                 low                        high
        ┌─────────────────────┬─────────────────────┐
  high  │ clicks on the       │ ★ add-to-cart rate  │
        │ hero button         │   checkout starts   │
sensi-  │                     │   step completion   │
tivity  │ "engagement"        │                     │
        ├─────────────────────┼─────────────────────┤
        │ scroll depth        │ revenue per visitor │
  low   │ time on page        │ lifetime value      │
        │                     │ retention at 90d    │
        └─────────────────────┴─────────────────────┘
                                        ↑
                          what everyone asks for, and
                          what almost nobody can power

The top-right cell is where testable metrics live. It’s a small box, and finding what belongs in it for your site is most of the work.

Why revenue per visitor is so expensive

Sensitivity is driven by variance relative to the mean — the coefficient of variation, the standard deviation divided by the mean. Revenue per visitor has an enormous one, because most visitors spend nothing and a few spend a great deal.

CONVERSION RATE                     REVENUE PER VISITOR

values: 0 or 1                      values: 0, 0, 0, 0, £34, 0, 0, £890, 0…
mean:   0.03                        mean:  £2.40
sd:     √(0.03 × 0.97) = 0.171      sd:    ~£28  (one order in a long tail)
CV:     0.171 / 0.03 = 5.7          CV:    28 / 2.40 = 11.7

                                    ~2× the CV → ~4× the sample,
                                    because sample scales with CV²

In plain terms: to detect the same relative change, revenue per visitor typically needs somewhere around four times the traffic of conversion rate — sometimes far more, because a single large order can move the mean of an arm on its own. That’s the whole reason the metric everyone wants is the metric nobody powers — Metric Sensitivity, Skewed and Heavy-Tailed Distributions.

Worked, with real numbers. 3% baseline conversion, detecting a 5% relative lift at 80% power and 95% confidence:

conversion rate        ≈ 210,000 visitors per arm
revenue per visitor    ≈ 840,000 per arm at 4× — more if the tail is heavy

at 40,000 visitors/week:   conversion      → 10.5 weeks
                           revenue/visitor → 42 weeks

Forty-two weeks is not a test. It’s a year in which the site, the traffic mix and the product have all changed — Test Duration.

Making the expensive metric affordable

Before giving up on it, three levers that genuinely work:

  • Winsorisation and Capping. Cap order values at, say, the 99th percentile. One £8,000 trade order shouldn’t decide a homepage test. This routinely cuts the variance substantially and biases the estimate only in the tail you capped
  • Variance Reduction, usually CUPED — controlled experiment using pre-experiment data. Subtract each user’s pre-test spend from their outcome. With a pre-period correlation of ρ = 0.5 the variance falls by ρ² = 25%, which is a 25% reduction in required sample for free
  • A ratio with a better denominator. Revenue per purchaser has far lower variance than revenue per visitor — but it’s conditioned on an outcome the test may have changed, so it can’t be a primary metric on its own — see the trap below

The surrogate trap

The standard move when the real metric is unaffordable is to test a closer, cheaper one — add-to-cart instead of purchase. It’s usually right, and it has a specific failure mode.

A surrogate is only valid if the treatment can’t move the surrogate without moving the outcome in the same direction. That condition fails constantly:

change: a large, prominent "Add to basket" button on the listing page

add-to-cart rate      +18%   ← the surrogate says ship it
checkout starts        +2%
orders                 −1%   ← people added the wrong size, then abandoned
revenue per visitor    −3%

The surrogate moved because the change made adding easier, not because it made buying more likely. This is Goodhart’s law arriving inside a two-week test.

How to use a surrogate safely:

  • Validate it historically. Across your archive of past tests, does the surrogate’s movement predict the outcome’s? If you have twenty tests with both measured, that’s an answerable question — and answering it is one of the best uses of an Experiment Archive
  • Keep the real metric as a guardrail, underpowered but watched. You can’t detect a 2% loss, but you can detect a 15% collapse — Guardrail Metrics
  • Prefer a surrogate downstream of the change, not upstream. Checkout starts is a better surrogate than button clicks because more of the funnel has to survive

The rules

  • One primary metric, named before launch — Overall Evaluation Criterion
  • Pick it from the sensitivity/meaningfulness box, not from the org chart. “The CFO wants revenue” is not a power calculation
  • Ratio metrics need the right denominator and the right variance formula. Revenue per visitor is a ratio of two random quantities, and its standard error is not the naive one — Ratio Metrics
  • Never change the primary metric after launch. If the chosen one turns out to be unpowered, the test is inconclusive and the next one is designed better — Pre-Registration
  • Write down what result would make you ship, in the metric’s units, before you see any data. If no plausible result changes the decision, the metric is wrong or the test is unnecessary

Where it interacts