Tags: statistics experimentation concept
Sample Size Calculation
Date: 2026-08-16
Four inputs, one output. Every term in the formula is a decision you’ve already made — the calculation just prices them, and the price is usually higher than anyone expects.
Sample size calculation works out how many units each arm of a test needs before it starts, so that an effect of the size you care about would be detected reliably if it’s there.
The formula, for a proportion
Per arm. Each piece:
| Term | Is | From |
|---|---|---|
| 1.96 at 95% significance | Your α — how often you’ll accept a false positive | |
| 0.84 at 80% power | Your power — how often you’ll accept a miss | |
| baseline rate, roughly | Your data. Get this right; everything scales off it | |
| the absolute difference you want to catch | Your MDE, converted from relative to absolute | |
| the variance of a proportion, doubled for two arms | Falls out of the maths, not a choice |
Worked
3% baseline, detecting a 10% relative lift, 80% power, 95% significance:
p₁ = 0.030 baseline
p₂ = 0.033 baseline × 1.10
p̄ = 0.0315 average
δ = 0.033 − 0.030 = 0.003 absolute difference
(1.96 + 0.84)² = 2.80² = 7.84
2 × 0.0315 × 0.9685 = 0.06102
δ² = 0.003² = 0.000009
n = 7.84 × 0.06102 ÷ 0.000009 = 53,150 per arm
53,000 per arm. 106,000 total. At 12,000 visitors a week, 106,000 ÷ 12,000 ≈ nine weeks.
In plain terms: if the change really does lift conversion from 3% to 3.3%, running 53,000 people through each version gives you an 80% chance of the test noticing — and if the change does nothing, only a 1-in-20 chance of the test wrongly claiming it did.
Reading the structure
Three things follow directly from where terms sit in the formula, and they’re worth more than the arithmetic:
- δ is squared, everything else isn’t. So the MDE dominates. Halving it costs 4×; the difference between 80% and 90% power costs only 1.35×. Spend your attention on the effect size you’re targeting, not on power
- enters through the variance term, so a lower baseline needs more traffic for the same relative lift. A 1% converting site needs roughly three times a 3% site to detect the same relative change. Low-converting sites are much harder to test on than their owners expect
- n is per arm. Three variants isn’t 1.5× the traffic, it’s 1.5× and a multiplicity correction that pushes it further
For continuous metrics
Revenue per visitor, order value, items per basket — replace the proportion variance with the actual variance:
Same example, revenue per visitor with mean £1.50 and standard deviation £13.44:
σ² = 0.03 × (60² + 50²) − 1.50² = 180.75 (£50 orders, SD £60)
δ = 10% × £1.50 = £0.15
n = 2 × 7.84 × 180.75 ÷ (0.15)² = 125,960 ≈ 126,000 per arm
2.4× the conversion-rate requirement for the same relative lift, and all of it is order-value spread. See Metric Sensitivity and Skewed and Heavy-Tailed Distributions.
Why calculators disagree
Run the same inputs through three tools and you’ll get three answers within a few percent of each other. Legitimate reasons:
- Pooled versus unpooled variance — whether or the two rates separately are used in the variance term
- Continuity correction — an adjustment for using a continuous distribution to approximate a discrete one
- One- versus two-tailed — halves the requirement, and is usually the wrong choice (One-Tailed vs Two-Tailed Tests)
- Relative versus absolute MDE input — the most common source of a wildly wrong answer, because the field is often unlabelled
Treat the output as a planning figure. A difference between 53,000 and 55,000 changes nothing; a difference between 53,000 and 530,000 means you entered the MDE in the wrong units.
Practical rules
- Calculate before building. It takes ten minutes and frequently ends the conversation
- Convert to a duration immediately. The sample size means nothing until it’s weeks, because weeks is what gets vetoed
- Never re-run it mid-test to justify stopping. That’s Peeking wearing a lab coat
- Round up and add a buffer for bots, internal traffic and assignment loss — see Bot and Internal Traffic
- Baseline from a full business cycle, not last week. A baseline measured over a promotional period sets an MDE you’ll never hit in normal trading
Assembled end to end in Guide - Statistics for CRO.