Tags: statistics experimentation guide
Guide - Statistics for CRO
Date: 2026-08-16
One test, start to finish, with the arithmetic done properly. Every concept links out — this guide does the working, not the explaining.
Runs on one example throughout: a UK retail site, 3% conversion rate, £50 average order value, 12,000 visitors a week.
1. Can you detect anything at all?
The first question, and the one most programmes skip. Answering it kills roughly half of proposed tests before they waste a month.
Get the baseline right first. Conversion rate over what denominator — users, sessions, or visitors? Session-scoped rates move when a timeout changes, so this decision is not cosmetic. See Sessionisation and Metric Design. Use the same denominator the test will use, measured over a full business cycle.
Then decide what’s worth finding. The Minimum Detectable Effect is a commercial decision, not a statistical one: what’s the smallest lift that would justify building this? Deciding it after seeing the data is P-Hacking.
The sample size calculation
For a conversion rate — a proportion — with two equal arms:
Detecting a 10% relative lift on a 3% baseline, at 80% power and 95% significance:
| Term | Value | Where from |
|---|---|---|
| 0.030 | baseline | |
| 0.033 | baseline + 10% relative | |
| 0.0315 | average of the two | |
| 1.96 | 95% significance, two-tailed | |
| 0.84 | 80% [[Statistical Power |
Working, step by step:
(1.96 + 0.84)² = 2.80² = 7.84
2 × 0.0315 × 0.9685 = 0.06102
(0.033 − 0.030)² = 0.003² = 0.000009
n = 7.84 × 0.06102 ÷ 0.000009 = 53,150 per arm
≈53,000 per arm, so ≈106,000 visitors in total.
Calculators disagree by a few percent depending on whether they pool the variance and whether they apply a continuity correction. Treat the output as a planning figure, not a threshold.
The duration check
106,000 ÷ 12,000 per week = 8.8 weeks. Call it nine.
That’s the number that kills tests. Nine weeks means one test per quarter on this page, no seasonal comparability, and a real chance the site changes underneath it. See Test Duration.
Your options at this point, in order of honesty:
- Test something bigger. A 10% MDE needing nine weeks means a 20% MDE needs about two and a quarter — sample size scales with the square of the effect, so doubling the MDE quarters the traffic
- Widen the scope — test site-wide rather than on one template, if the change makes sense there
- Reduce the variance — see §7
- Don’t test it. Ship it on judgement, or don’t build it. An underpowered test is worse than no test, because it produces a number people will act on
Minimum duration regardless of sample size: one full week, and ideally two. Weekday and weekend traffic convert differently, and a test that hits its sample in three days has sampled three days of behaviour, not your customers.
2. Metric choice decides the traffic bill
The same site, the same MDE, a different metric:
| Metric | Standard deviation | n per arm |
|---|---|---|
| Conversion rate | — (binary) | 53,000 |
| Revenue per visitor | £13.44 | 126,000 |
Revenue per visitor working. With 3% converting at a mean order of £50 and an order-value standard deviation of £60:
E[X] = 0.03 × 50 = £1.50
E[X²] = 0.03 × (50² + 60²) = 0.03 × 6100 = 183
Var = 183 − 1.50² = 180.75
σ = √180.75 = £13.44
n = 2 × 7.84 × 180.75 ÷ (0.15)² = 126,000 per arm
2.4× the traffic for the same relative lift, and the entire difference is order-value spread — see Skewed and Heavy-Tailed Distributions. Revenue per visitor is the metric everyone wants and few sites can afford.
The practical rule: test on the most sensitive metric that still answers the question. Conversion rate as primary, revenue per visitor as a secondary you report with a wide interval and don’t make decisions on. See Metric Sensitivity and Overall Evaluation Criterion.
Ratio metrics deserve care — the variance of a ratio isn’t the variance of its parts, and tools get this wrong. See Ratio Metrics.
3. Which test
For a conversion rate comparison between two groups: a two-proportion test. Chi-squared and the z-test for proportions give the same answer here; they’re the same question asked twice.
That covers nearly everything you’ll run. The exceptions, and what they need, are in Guide - Choosing a Statistical Test. If your metric is continuous and heavy-tailed, Bootstrapping is usually more honest than assuming a distribution.
4. While it runs
Don’t look. This is the single highest-value discipline in the guide, and the most commonly broken.
A fixed-horizon test is designed so that the false positive rate is 5% when you check once, at the end. Every additional look is another chance to cross the threshold by luck. Five looks, treated naively as independent:
P(at least one false positive) = 1 − 0.95⁵ = 22.6%
Looks are correlated rather than independent, so the true figure is lower — but the direction and rough scale hold. You designed a 5% error rate and you’re running at something closer to 15%. See Peeking.
In plain terms: stopping a test the moment it looks significant means you have selected for the moment it looked good. That moment happens by chance in tests where nothing is happening at all.
If you genuinely need to monitor continuously, use a design built for it — Sequential Testing or Always-Valid Inference. Those spend some power to buy the licence to look. Retrofitting the licence onto a fixed-horizon test is not available.
Do check these, daily, without looking at the result:
- Sample Ratio Mismatch — a 50/50 split arriving as 52/48 on large numbers invalidates the test regardless of what it says. It means assignment is broken, and no analysis survives it
- Guardrails — page speed, error rates, add-to-cart. Stop for harm, never for success. See Guardrail Metrics
- Both variants still render — the most common failed experiment is one that stopped showing the variant
5. Reading the result
In this order. Skipping to the last line is how bad decisions get made.
- SRM check — failed, stop. There is nothing to read
- Did it run its planned duration and sample? Short, and the p-value doesn’t mean what it says
- Guardrails — anything broken, the win is irrelevant
- Primary metric only. The one nominated before launch
- Everything else — diagnostic, never decisive
Significance is not the answer
Suppose the primary metric comes back:
- Control 3.00%, variant 3.09% — a +3.0% relative lift
- p = 0.03, so significant at 95%
- 95% confidence interval on the relative lift: +0.3% to +5.7%
The tempting reading is “we made 3%”. The honest reading is the interval. On 500,000 annual visitors at £50:
| Scenario | Extra orders/year | Extra revenue |
|---|---|---|
| CI lower bound (+0.3%) | 45 | £2,250 |
| Point estimate (+3.0%) | 450 | £22,500 |
| CI upper bound (+5.7%) | 855 | £42,750 |
The result is “somewhere between £2,250 and £42,750 a year”. If the build cost £15,000, this test did not tell you to ship. It told you that you still don’t know — which is a real finding, and much more useful than a false certainty.
That gap between statistically detectable and commercially worthwhile is Practical vs Statistical Significance, and it’s how a testing programme passes every check for a year while producing nothing.
When it isn’t significant
Not significant means “we didn’t detect an effect”, never “there is no effect.” With 80% power you would miss a real effect of exactly your MDE one time in five, by design.
Useful response: look at the confidence interval. If it spans −1% to +8%, you learned nothing. If it spans −0.4% to +0.6%, you’ve learned there’s no large effect here, which is worth knowing. See Inconclusive Results.
6. How results turn out to be wrong
Ranked by how often they actually bite.
- You checked early and stopped. §4
- You tested several metrics and reported the one that won. Twenty metrics at 95% confidence produces one false positive by construction — The Multiple Comparisons Problem
- You segmented afterwards and found a winner in mobile. Post-hoc slicing generates significance from noise reliably. If a segment matters, it was named before launch — Segmentation (test results)
- The shipped change underdelivered. Expected. You selected the winner on a noisy estimate, so the estimate was biased high the moment you chose it — Winner’s Curse
- The effect faded. Three candidates: Novelty and Primacy Effects if regulars reacted to the change itself, Regression to the Mean if you intervened on something already at an extreme, or winner’s curse above
- Every segment improved but the total didn’t. Simpson’s Paradox — traffic mix moved
7. Getting more from the same traffic
Before concluding a test is unaffordable:
-
Winsorisation and Capping — cap order values at, say, the 99th percentile. In the §2 example, cutting the order-value standard deviation from £60 to £40 takes the required sample from 126,000 to about 84,000 per arm, a third less traffic:
E[X²] = 0.03 × (50² + 40²) = 0.03 × 4100 = 123 Var = 123 − 1.50² = 120.75 n = 2 × 7.84 × 120.75 ÷ (0.15)² = 84,150Indicative rather than exact — capping also pulls the mean down slightly, which shrinks the absolute effect you’re detecting. The cap must be set before you look at results, and applied identically to both arms
-
Variance Reduction — CUPED uses pre-experiment behaviour as a covariate, and can cut required sample substantially where users have history to draw on [CHECK: typical reduction range against a current source before quoting a figure]
-
Move to a more sensitive primary metric — §2
-
Widen the exposure — testing site-wide rather than one template is often the cheapest way to find the traffic
-
Increase the MDE — the square relationship means this is powerful. Be honest that you’re now testing a different question
8. Saying it out loud
The last step, and the one that determines whether the work mattered. See Communicating Uncertainty.
A defensible result statement has four parts:
“The variant increased conversion rate by 3.0% relative (point estimate), with a 95% confidence interval of +0.3% to +5.7% (uncertainty). At current traffic that’s £2,000 to £43,000 a year (commercial translation). Given a £15,000 build cost, I’d re-run it with more traffic before committing (recommendation).”
Things not to say, and why:
| Don’t | Because |
|---|---|
| ”We’re 95% confident the true lift is 3%“ | That isn’t what a confidence interval says |
| ”The test proved it works” | Nothing here proves anything |
| ”It wasn’t significant, so it doesn’t work” | See §5 |
| ”Mobile users converted 12% better” | Post-hoc segment, unless pre-registered |
The short version
Five things, if you retain nothing else:
- Calculate the sample size before building anything. Most tests are unaffordable and it takes ten minutes to find out
- Nominate one primary metric, before launch, in writing
- Don’t look until it’s finished
- Read the confidence interval, not the p-value. Then convert it to pounds
- “Not significant” is not “no effect”
Full domain: Statistics MOC. The experiment procedure around this: Guide - Running an Experiment.