Tags: experimentation statistics concept

Winner’s Curse

Date: 2026-08-16


The tests you ship systematically overstate their effect — not because anyone cheated, but because you selected them for having crossed a line, and crossing a line is easier with a favourable roll. The bias is worst in exactly the underpowered tests that need it least.


What it is

The winner’s curse is the upward bias in the measured effect of any result selected for being large or significant.

Selection is the whole mechanism. Among tests that crossed the significance threshold, those that got a lucky sample are over-represented, because luck helped them cross. The average measured effect of the winners is therefore higher than the average true effect of the same changes.

Worked

Ten tests, each with a true effect of exactly +2% relative. The design has 50% power to detect that effect, so roughly half will reach significance.

                true     observed    significant?
test 1          +2%        +5.1%          ✓        ← lucky
test 2          +2%        +4.4%          ✓        ← lucky
test 3          +2%        +3.9%          ✓        ← lucky
test 4          +2%        +3.6%          ✓        ← lucky
test 5          +2%        +3.4%          ✓        ← lucky
test 6          +2%        +1.2%          ✗
test 7          +2%        +0.6%          ✗
test 8          +2%        +0.1%          ✗
test 9          +2%        −0.8%          ✗
test 10         +2%        −1.4%          ✗
                        ────────
mean of all ten            +2.0%   ← unbiased
mean of the winners        +4.1%   ← what you ship, 2× the truth

Nothing is wrong with any individual test. Each measured its sample correctly. The bias appears only at the moment you filter to the winners — which is the moment you decide what to ship.

In plain terms: you picked the results that looked best. Looking best requires being lucky as well as being real, so the ones you picked were luckier than average, and luck doesn’t repeat.

Power is the lever

The size of the bias depends almost entirely on Statistical Power:

PowerInflation of shipped winners
90%Small — most real effects cross without luck
80%Modest
50%Roughly 2×
20%Severe — winners can be several times their true effect

At low power, the only way an effect crosses the line is with substantial help from noise, so every winner is inflated. This is why underpowered testing doesn’t just miss things — it actively produces exaggerated wins, and a programme running at 30% power will report a strong year and deliver nothing.

What it explains

  • Shipped changes underdelivering. The single most common complaint about testing programmes, and it’s expected behaviour rather than evidence the test was wrong
  • Forecasts built from test results overshooting. Summing the measured lifts of a year’s winners and projecting them forward overstates by the inflation factor, compounding
  • Segment findings being wildly inflated. A segment slice is a small sample, so it’s low-powered, so its winners are the most exaggerated of all — Segmentation (test results)
  • The best-looking test in a batch being the least replicable. Selecting the largest effect from a set is selecting for luck twice over

What to do

  • Power tests properly. The only real fix. 80% minimum, higher for anything expensive — Sample Size Calculation
  • Forecast from the confidence interval’s lower bound, not the point estimate. Deliberately conservative, and closer to what you’ll get — Confidence Intervals
  • Post-Test Validation. Measure the shipped change again in production. Expect regression towards the true effect; a fall of a third to a half is normal at 80% power, not a sign of a broken test
  • Don’t rank ideas by past measured lift when prioritising. The biggest measured winners are the most inflated
  • Discount when aggregating. A programme reporting “£400,000 of tested wins this year” should expect materially less to appear in the accounts, and saying so in advance protects credibility

Regression to the Mean is the same statistical phenomenon applied to units selected on a prior measurement — the worst-performing pages improving without intervention. Winner’s curse is that mechanism applied to test results selected on significance. Same maths, different thing being selected.