Tags: statistics experimentation concept
Statistical Power
Date: 2026-08-16
The probability of detecting an effect of a given size, if it’s really there. At the conventional 80%, you miss one real effect in five — and an underpowered test that comes back flat tells you nothing at all, while looking exactly like one that does.
What it is
Power = 1 − β, where β is the Type II error rate. It’s always power to detect an effect of a specific size — there’s no such thing as the power of a test in the abstract.
80% power to detect a 10% relative lift means: if a 10% lift is genuinely there, this design finds it four times out of five. Smaller lifts, it finds less often. That’s not a defect; it’s the design working.
The four levers
Power rises with:
| Lever | Direction | Practical reality |
|---|---|---|
| Sample size | more → more power | The main lever. Slow and expensive |
| Effect size | bigger → more power | Not yours to choose, but the [[Minimum Detectable Effect |
| Significance level α | looser → more power | Buys power by accepting more false positives. Rarely the right trade |
| Variance | lower → more power | The underrated one — see Variance Reduction and Winsorisation and Capping |
Nothing else. If a test lacks power, one of those four has to move, and “run it a bit longer and see” is the sample size lever applied without arithmetic.
Worked: the running example
3% baseline, 12,000 visitors a week, detecting a 10% relative lift at 95% significance:
| Power | n per arm | Total | Weeks |
|---|---|---|---|
| 50% | 26,000 | 52,000 | 4.3 |
| 80% | 53,000 | 106,000 | 8.8 |
| 90% | 71,000 | 142,000 | 11.8 |
| 95% | 88,000 | 176,000 | 14.7 |
Going from 80% to 95% power costs another six weeks — a two-thirds increase in traffic to reduce your miss rate from one-in-five to one-in-twenty. Whether that’s worth it depends on what missing costs.
The 50% row is worth sitting with: a design at 50% power is a coin toss on whether you see a real effect. Plenty of real tests run at roughly this, because nobody calculated.
Why underpowered testing is worse than not testing
Three compounding problems, and only the first is obvious.
- You miss real wins and never learn that you did. The test says “no significant difference” and the idea is dead. Nothing distinguishes that output from a genuinely flat result — see Inconclusive Results
- The wins you do find are exaggerated. In an underpowered design, the only effects large enough to cross the significance line are ones that got a favourable roll. So your significant results systematically overstate the truth, and shipping them underdelivers — the Winner’s Curse
- The programme learns the wrong lesson. A year of underpowered tests produces a low win rate and a few inflated winners, which reads as “testing doesn’t work here” rather than “we never had the traffic” — Win Rate and Expected Value
In plain terms: a test without enough traffic doesn’t give you a weaker answer. It gives you an answer that’s wrong in a specific direction — invisible misses, plus overstated wins.
Post-hoc power is meaningless
Calculating power after the test, using the observed effect, is a common and empty ritual. It’s a deterministic re-expression of the p-value — a non-significant result always yields low observed power, so it can only ever tell you what you already knew.
The useful question after a flat result isn’t “was I powered?” but “what effects can I now rule out?” — which is the confidence interval. An interval of −0.4% to +0.6% rules out anything large. An interval of −4% to +8% rules out nothing, and that’s your answer about power.
In practice
- Calculate before building. Power analysis takes ten minutes and kills roughly half of proposed tests before they cost anything — Guide - Statistics for CRO
- 80% is a convention, not a requirement. Consider higher for expensive or irreversible changes, lower where you’ll retest anyway
- Report the MDE alongside a null result. “No significant difference; this test could only have detected a lift above 9%” is honest and actionable. “No significant difference” alone is not
- Guardrails need their own power calculation. A guardrail that can’t detect the harm you’re worried about is decoration