Tags: statistics experimentation concept
Variance Reduction
Date: 2026-08-16
Use what you already knew about a user before the test started to subtract predictable variation from their outcome. Same traffic, same effect, narrower interval — the only lever that buys power without costing anything.
What it is
Variance reduction is any technique that lowers the noise in a metric without changing what the metric measures, so the same traffic produces a narrower interval.
It doesn’t make effects bigger or results more favourable. It makes the measurement less blurry.
The idea
Sample size scales with variance. Most of the variance in a user’s outcome isn’t caused by your test at all — it’s caused by the user. Heavy buyers buy heavily in both arms.
If you can predict part of someone’s outcome from data that predates the experiment, you can subtract that prediction and analyse only the residual. The treatment effect survives untouched; the noise shrinks.
Crucially the covariate must be pre-experiment. Anything measured during the test can be affected by the treatment, and adjusting for it would remove part of the effect you’re measuring — or manufacture one.
CUPED
CUPED — controlled experiment using pre-experiment data — is the standard implementation.
For each user, take their outcome Y and a pre-period covariate X, usually the same metric measured over the weeks before the test:
where θ is chosen to maximally cancel the shared variation — the covariance of X and Y divided by the variance of X.
The variance that remains is:
with ρ the correlation between pre-period and in-test behaviour. The whole method reduces to that one number.
ρ = 0.3 → variance × 0.91 → 9% less sample needed
ρ = 0.5 → variance × 0.75 → 25% less
ρ = 0.7 → variance × 0.51 → 49% less
Worked, on the running example. Revenue per visitor needs 126,000 per arm. With a pre-period correlation of 0.5:
variance multiplier = 1 − 0.5² = 0.75
n = 126,000 × 0.75 = 94,500 per arm
before 126,000 × 2 arms ÷ 12,000/week = 21 weeks
after 94,500 × 2 arms ÷ 12,000/week = 15.75 ≈ 16 weeks
Five weeks saved, from data you already had.
In plain terms: people who spent a lot last month will mostly spend a lot this month, whichever version they see. Account for that and what’s left is a much clearer view of what the test itself changed — the same answer, found sooner.
What ρ you’ll actually get
The honest constraint, and the reason CUPED helps some sites and not others.
| Situation | Typical correlation | Worth it? |
|---|---|---|
| Logged-in users with months of history | Moderate to strong | Yes — the main use case |
| Returning identified customers | Moderate | Usually |
| Mostly-anonymous ecommerce traffic | Weak to none | Rarely |
| First-time visitors | Zero by definition | No |
This is the catch for retail CRO. If most of your traffic is anonymous and non-returning, there is no pre-period to draw on, and the technique has nothing to work with. Its natural home is subscription and product analytics, where users are identified and have history. How much of your traffic qualifies depends directly on Identity Stitching.
[CHECK: published typical variance reductions from CUPED — quote a source rather than a remembered range before putting a figure in a report.]
Other forms
- Stratification — split assignment by a pre-known attribute (device, new/returning, country) so each stratum is balanced, then combine within-stratum effects. Simpler than CUPED, no modelling, and it also removes the Simpson’s Paradox mix risk
- Regression adjustment — include pre-period covariates in a regression of outcome on treatment. Mathematically close to CUPED, more flexible, easier to misuse
- Better metric choice — a lower-variance primary is variance reduction by another name, and it’s free — Metric Sensitivity
- Winsorisation and Capping — attacks the same problem from the tail rather than the covariate
Rules
- Covariate strictly pre-experiment. The single hard rule. Using in-test data biases the estimate
- Choose the covariate before launch, and record it. Trying several and keeping the one that produces significance is P-Hacking with extra steps
- Validate on an A test. CUPED applied to an A/A test should produce a narrower interval around zero, not a significant result. If it produces significance, the implementation is wrong
- It does not fix bias. Variance reduction narrows the interval around whatever you were measuring. If assignment is broken, it makes a wrong answer look more precise — check Sample Ratio Mismatch first