Tags: statistics experimentation concept

Variance Reduction

Date: 2026-08-16


Use what you already knew about a user before the test started to subtract predictable variation from their outcome. Same traffic, same effect, narrower interval — the only lever that buys power without costing anything.


What it is

Variance reduction is any technique that lowers the noise in a metric without changing what the metric measures, so the same traffic produces a narrower interval.

It doesn’t make effects bigger or results more favourable. It makes the measurement less blurry.

The idea

Sample size scales with variance. Most of the variance in a user’s outcome isn’t caused by your test at all — it’s caused by the user. Heavy buyers buy heavily in both arms.

If you can predict part of someone’s outcome from data that predates the experiment, you can subtract that prediction and analyse only the residual. The treatment effect survives untouched; the noise shrinks.

Crucially the covariate must be pre-experiment. Anything measured during the test can be affected by the treatment, and adjusting for it would remove part of the effect you’re measuring — or manufacture one.

CUPED

CUPED — controlled experiment using pre-experiment data — is the standard implementation.

For each user, take their outcome Y and a pre-period covariate X, usually the same metric measured over the weeks before the test:

where θ is chosen to maximally cancel the shared variation — the covariance of X and Y divided by the variance of X.

The variance that remains is:

with ρ the correlation between pre-period and in-test behaviour. The whole method reduces to that one number.

ρ = 0.3    →  variance × 0.91   →   9% less sample needed
ρ = 0.5    →  variance × 0.75   →  25% less
ρ = 0.7    →  variance × 0.51   →  49% less

Worked, on the running example. Revenue per visitor needs 126,000 per arm. With a pre-period correlation of 0.5:

variance multiplier = 1 − 0.5²  = 0.75
n = 126,000 × 0.75              = 94,500 per arm
before   126,000 × 2 arms ÷ 12,000/week  = 21 weeks
after     94,500 × 2 arms ÷ 12,000/week  = 15.75 ≈ 16 weeks

Five weeks saved, from data you already had.

In plain terms: people who spent a lot last month will mostly spend a lot this month, whichever version they see. Account for that and what’s left is a much clearer view of what the test itself changed — the same answer, found sooner.

What ρ you’ll actually get

The honest constraint, and the reason CUPED helps some sites and not others.

SituationTypical correlationWorth it?
Logged-in users with months of historyModerate to strongYes — the main use case
Returning identified customersModerateUsually
Mostly-anonymous ecommerce trafficWeak to noneRarely
First-time visitorsZero by definitionNo

This is the catch for retail CRO. If most of your traffic is anonymous and non-returning, there is no pre-period to draw on, and the technique has nothing to work with. Its natural home is subscription and product analytics, where users are identified and have history. How much of your traffic qualifies depends directly on Identity Stitching.

[CHECK: published typical variance reductions from CUPED — quote a source rather than a remembered range before putting a figure in a report.]

Other forms

  • Stratification — split assignment by a pre-known attribute (device, new/returning, country) so each stratum is balanced, then combine within-stratum effects. Simpler than CUPED, no modelling, and it also removes the Simpson’s Paradox mix risk
  • Regression adjustment — include pre-period covariates in a regression of outcome on treatment. Mathematically close to CUPED, more flexible, easier to misuse
  • Better metric choice — a lower-variance primary is variance reduction by another name, and it’s free — Metric Sensitivity
  • Winsorisation and Capping — attacks the same problem from the tail rather than the covariate

Rules

  • Covariate strictly pre-experiment. The single hard rule. Using in-test data biases the estimate
  • Choose the covariate before launch, and record it. Trying several and keeping the one that produces significance is P-Hacking with extra steps
  • Validate on an A test. CUPED applied to an A/A test should produce a narrower interval around zero, not a significant result. If it produces significance, the implementation is wrong
  • It does not fix bias. Variance reduction narrows the interval around whatever you were measuring. If assignment is broken, it makes a wrong answer look more precise — check Sample Ratio Mismatch first