Tags: statistics experimentation concept
Minimum Detectable Effect
Date: 2026-08-16
The smallest effect your design can reliably find. It’s a commercial decision disguised as a statistical input — and because sample size scales with its square, halving it quadruples your traffic bill.
What it is
Minimum detectable effect (MDE) is the smallest true effect a test would detect, at your chosen power and significance. Effects smaller than it will usually be missed; that’s the design working as specified.
It’s an input, not an output. You choose it, then the maths tells you what it costs.
The square relationship
The single most useful fact in test planning. Sample size is inversely proportional to the square of the effect:
On the running example — 3% baseline, 12,000 visitors a week, 80% power, 95% significance:
| MDE (relative) | Absolute lift | n per arm | Total | Weeks |
|---|---|---|---|---|
| 2% | 3.00% → 3.06% | 1,281,000 | 2,562,000 | 214 |
| 5% | 3.00% → 3.15% | 208,000 | 416,000 | 35 |
| 10% | 3.00% → 3.30% | 53,000 | 106,000 | 8.8 |
| 20% | 3.00% → 3.60% | 13,900 | 27,800 | 2.3 |
| 30% | 3.00% → 3.90% | 6,500 | 13,000 | 1.1 |
Read the 5% and 10% rows together. Halving the MDE took the test from nine weeks to thirty-five — four times the traffic for twice the sensitivity. Going to 2% is four years, which is a polite way of saying it cannot be tested.
This is why “let’s just run it longer” fails as a strategy. Doubling duration doesn’t halve your MDE, it improves it by about 30%.
Choosing it
The mistake is picking the number you hope for. The MDE should come from the decision:
What’s the smallest lift that would make this worth having? Include the build cost, the maintenance, and the opportunity cost of the test slot.
Worked: a change costing £15,000 to build, on 500,000 annual visitors at £50 order value.
break-even extra revenue = £15,000
orders needed = 15,000 ÷ 50 = 300
baseline annual orders = 500,000 × 0.03 = 15,000
required relative lift = 300 ÷ 15,000 = 2%
Break-even is a 2% lift — which the table says is untestable. That’s a genuinely useful finding: this change cannot be validated by experiment. Ship it on judgement, or don’t build it, but don’t run a test that can’t answer the question.
If you’d instead want it to pay back three times over in year one, the MDE is 6%, which is roughly 25 weeks. Still probably not viable.
What to do when the MDE is unaffordable
In rough order of preference:
- Test somewhere with more traffic. Site-wide rather than one template; the homepage rather than a category page
- Test a bolder change. A 20% MDE is two weeks. Timid changes are expensive precisely because they’re timid — the traffic cost of detecting them is enormous
- Reduce variance — Variance Reduction, Winsorisation and Capping, or a more sensitive primary metric (Metric Sensitivity)
- Accept lower power for cheap reversible changes. 60% power halves the sample against 80%, at the cost of missing two real effects in five
- Don’t test. Ship on judgement with a monitoring plan, or use a holdout to measure the cumulative effect of many such changes at once
In plain terms: the MDE is the site telling you what questions it’s big enough to answer. Most sites can answer “did this big change help?” and cannot answer “did this small change help?” — and no amount of patience converts one into the other.
Failure modes
- Setting the MDE after seeing the result — retroactively deciding you were looking for the effect you found. P-Hacking
- Quoting a null result without its MDE. “No difference” from a test that could only detect a 15% lift is not evidence of no difference
- Confusing MDE with the expected effect. The MDE is a floor on detectability, not a prediction. Most real effects are smaller than the MDE people pick
- Relative and absolute confusion — a 10% relative lift on a 3% baseline is 0.3 percentage points. Mixing them up misprices tests by an order of magnitude. See Effect Size