Tags: statistics experimentation concept

Metric Sensitivity

Date: 2026-08-16


How much traffic a metric needs to detect a given change. It’s driven by variance relative to the mean, which is why revenue per visitor — the metric everyone wants — is the one most sites can’t afford to test on.


What it is

Metric sensitivity is how readily a metric reveals a change of a given size — in practice, how much traffic it needs before an effect becomes detectable.

A sensitive metric moves visibly on modest effects. An insensitive one can be genuinely improving while every test on it comes back inconclusive.

What drives it

Sample size scales with variance. The useful summary of variance relative to size is the coefficient of variation — the standard deviation divided by the mean:

Higher CV means a noisier metric means more traffic for the same MDE. Everything below follows from that one relationship.

The ranking

Running example: 3% conversion, £50 mean order value with £60 standard deviation, detecting a 10% relative lift at 80% power.

MetricMeanσCVn per arm
Add-to-cart rate12%0.322.712,000
Conversion rate3%0.175.753,000
Revenue per visitor£1.50£13.449.0126,000
AOV (converters only)£50£601.22,300 converters ≈ 75,000 visitors

In plain terms: the noisier a number is from visitor to visitor, the more visitors you need before a real change stands out from the random bouncing around. Revenue per visitor is mostly zeros with the occasional large order, so it bounces a lot.

Two things worth reading off it.

Rates higher up the funnel are far cheaper. Add-to-cart at 12% needs a quarter of the traffic that conversion at 3% does, because a proportion’s variance is p(1−p) — highest at 50% in absolute terms, but relative to the mean it falls fast as p rises. Low-converting sites are structurally expensive to test on, which is a hard thing to explain to whoever owns a 0.8% converting site.

AOV looks cheap and isn’t. Its CV is the lowest on the table by far, and it needs only 2,300 converters — but at 3% conversion, gathering 2,300 converters takes 75,000 visitors. Metrics measured on a subset get quoted in the wrong unit, and the conversion filter costs more than the low variance saves.

The trade you’re actually making

Sensitivity and meaningfulness point in opposite directions.

   add-to-cart rate  ────────────────────  revenue per visitor
   cheap to move                            what the business cares about
   cheap to measure                         expensive to measure
   might not translate                      translates by definition

Move up the funnel and you can detect changes but may be measuring something that doesn’t reach revenue. Move down and you’re measuring the right thing with too little power to see it.

The workable arrangement, and it’s a convention worth adopting:

  • Primary metric: the most sensitive one that still plausibly implies the business outcome. Usually conversion rate — Overall Evaluation Criterion
  • Secondary: revenue per visitor, reported with its interval, explicitly not decision-eligible
  • Diagnostic: upper-funnel rates, to explain why the primary moved

What this avoids is the common failure of nominating revenue per visitor as primary, running underpowered, and calling every test inconclusive.

Raising sensitivity without more traffic

  • Winsorisation and Capping — most of revenue per visitor’s CV is order-value spread. Capping the tail cuts it directly
  • Variance Reduction — CUPED (controlled experiment using pre-experiment data) and covariates, using pre-period behaviour
  • Pick a better denominator. User-scoped rather than session-scoped removes the noise that Sessionisation injects
  • Bound the exposure. Only counting users who reached the page under test raises the baseline rate, which raises sensitivity — but changes the question to “among people who got there”
  • Composite metrics — combining several signals into one score. Sensitive, and frequently uninterpretable, so use sparingly

Watch for

  • Quoting sample size in the wrong unit. “2,300 for AOV” is converters, not visitors, and at a 3% baseline the difference between the two is 33×
  • Metrics with a floor or ceiling. A metric already at 94% has little room, so relative lifts are arithmetically constrained
  • Very rare events. Sub-1% metrics are effectively untestable outside enormous traffic; measure them as guardrails rather than primaries
  • Sensitivity gains that change the question. Bounding exposure is legitimate; quietly redefining the metric mid-programme is Metric Drift