Tags: statistics experimentation concept

Sample Size Calculation

Date: 2026-08-16


Four inputs, one output. Every term in the formula is a decision you’ve already made — the calculation just prices them, and the price is usually higher than anyone expects.


Sample size calculation works out how many units each arm of a test needs before it starts, so that an effect of the size you care about would be detected reliably if it’s there.

The formula, for a proportion

Per arm. Each piece:

TermIsFrom
1.96 at 95% significanceYour α — how often you’ll accept a false positive
0.84 at 80% powerYour power — how often you’ll accept a miss
baseline rate, roughlyYour data. Get this right; everything scales off it
the absolute difference you want to catchYour MDE, converted from relative to absolute
the variance of a proportion, doubled for two armsFalls out of the maths, not a choice

Worked

3% baseline, detecting a 10% relative lift, 80% power, 95% significance:

p₁ = 0.030                       baseline
p₂ = 0.033                       baseline × 1.10
p̄  = 0.0315                      average
δ  = 0.033 − 0.030 = 0.003       absolute difference

(1.96 + 0.84)²        = 2.80²           = 7.84
2 × 0.0315 × 0.9685                     = 0.06102
δ²                    = 0.003²          = 0.000009

n = 7.84 × 0.06102 ÷ 0.000009           = 53,150 per arm

53,000 per arm. 106,000 total. At 12,000 visitors a week, 106,000 ÷ 12,000 ≈ nine weeks.

In plain terms: if the change really does lift conversion from 3% to 3.3%, running 53,000 people through each version gives you an 80% chance of the test noticing — and if the change does nothing, only a 1-in-20 chance of the test wrongly claiming it did.

Reading the structure

Three things follow directly from where terms sit in the formula, and they’re worth more than the arithmetic:

  • δ is squared, everything else isn’t. So the MDE dominates. Halving it costs 4×; the difference between 80% and 90% power costs only 1.35×. Spend your attention on the effect size you’re targeting, not on power
  • enters through the variance term, so a lower baseline needs more traffic for the same relative lift. A 1% converting site needs roughly three times a 3% site to detect the same relative change. Low-converting sites are much harder to test on than their owners expect
  • n is per arm. Three variants isn’t 1.5× the traffic, it’s 1.5× and a multiplicity correction that pushes it further

For continuous metrics

Revenue per visitor, order value, items per basket — replace the proportion variance with the actual variance:

Same example, revenue per visitor with mean £1.50 and standard deviation £13.44:

σ² = 0.03 × (60² + 50²) − 1.50²   = 180.75     (£50 orders, SD £60)
δ  = 10% × £1.50                 = £0.15
n  = 2 × 7.84 × 180.75 ÷ (0.15)²  = 125,960  ≈ 126,000 per arm

2.4× the conversion-rate requirement for the same relative lift, and all of it is order-value spread. See Metric Sensitivity and Skewed and Heavy-Tailed Distributions.

Why calculators disagree

Run the same inputs through three tools and you’ll get three answers within a few percent of each other. Legitimate reasons:

  • Pooled versus unpooled variance — whether or the two rates separately are used in the variance term
  • Continuity correction — an adjustment for using a continuous distribution to approximate a discrete one
  • One- versus two-tailed — halves the requirement, and is usually the wrong choice (One-Tailed vs Two-Tailed Tests)
  • Relative versus absolute MDE input — the most common source of a wildly wrong answer, because the field is often unlabelled

Treat the output as a planning figure. A difference between 53,000 and 55,000 changes nothing; a difference between 53,000 and 530,000 means you entered the MDE in the wrong units.

Practical rules

  • Calculate before building. It takes ten minutes and frequently ends the conversation
  • Convert to a duration immediately. The sample size means nothing until it’s weeks, because weeks is what gets vetoed
  • Never re-run it mid-test to justify stopping. That’s Peeking wearing a lab coat
  • Round up and add a buffer for bots, internal traffic and assignment loss — see Bot and Internal Traffic
  • Baseline from a full business cycle, not last week. A baseline measured over a promotional period sets an MDE you’ll never hit in normal trading

Assembled end to end in Guide - Statistics for CRO.