Tags: statistics experimentation concept

Statistical Power

Date: 2026-08-16


The probability of detecting an effect of a given size, if it’s really there. At the conventional 80%, you miss one real effect in five — and an underpowered test that comes back flat tells you nothing at all, while looking exactly like one that does.


What it is

Power = 1 − β, where β is the Type II error rate. It’s always power to detect an effect of a specific size — there’s no such thing as the power of a test in the abstract.

80% power to detect a 10% relative lift means: if a 10% lift is genuinely there, this design finds it four times out of five. Smaller lifts, it finds less often. That’s not a defect; it’s the design working.

The four levers

Power rises with:

LeverDirectionPractical reality
Sample sizemore → more powerThe main lever. Slow and expensive
Effect sizebigger → more powerNot yours to choose, but the [[Minimum Detectable Effect
Significance level αlooser → more powerBuys power by accepting more false positives. Rarely the right trade
Variancelower → more powerThe underrated one — see Variance Reduction and Winsorisation and Capping

Nothing else. If a test lacks power, one of those four has to move, and “run it a bit longer and see” is the sample size lever applied without arithmetic.

Worked: the running example

3% baseline, 12,000 visitors a week, detecting a 10% relative lift at 95% significance:

Powern per armTotalWeeks
50%26,00052,0004.3
80%53,000106,0008.8
90%71,000142,00011.8
95%88,000176,00014.7

Going from 80% to 95% power costs another six weeks — a two-thirds increase in traffic to reduce your miss rate from one-in-five to one-in-twenty. Whether that’s worth it depends on what missing costs.

The 50% row is worth sitting with: a design at 50% power is a coin toss on whether you see a real effect. Plenty of real tests run at roughly this, because nobody calculated.

Why underpowered testing is worse than not testing

Three compounding problems, and only the first is obvious.

  1. You miss real wins and never learn that you did. The test says “no significant difference” and the idea is dead. Nothing distinguishes that output from a genuinely flat result — see Inconclusive Results
  2. The wins you do find are exaggerated. In an underpowered design, the only effects large enough to cross the significance line are ones that got a favourable roll. So your significant results systematically overstate the truth, and shipping them underdelivers — the Winner’s Curse
  3. The programme learns the wrong lesson. A year of underpowered tests produces a low win rate and a few inflated winners, which reads as “testing doesn’t work here” rather than “we never had the traffic” — Win Rate and Expected Value

In plain terms: a test without enough traffic doesn’t give you a weaker answer. It gives you an answer that’s wrong in a specific direction — invisible misses, plus overstated wins.

Post-hoc power is meaningless

Calculating power after the test, using the observed effect, is a common and empty ritual. It’s a deterministic re-expression of the p-value — a non-significant result always yields low observed power, so it can only ever tell you what you already knew.

The useful question after a flat result isn’t “was I powered?” but “what effects can I now rule out?” — which is the confidence interval. An interval of −0.4% to +0.6% rules out anything large. An interval of −4% to +8% rules out nothing, and that’s your answer about power.

In practice

  • Calculate before building. Power analysis takes ten minutes and kills roughly half of proposed tests before they cost anything — Guide - Statistics for CRO
  • 80% is a convention, not a requirement. Consider higher for expensive or irreversible changes, lower where you’ll retest anyway
  • Report the MDE alongside a null result. “No significant difference; this test could only have detected a lift above 9%” is honest and actionable. “No significant difference” alone is not
  • Guardrails need their own power calculation. A guardrail that can’t detect the harm you’re worried about is decoration