Tags: experimentation statistics concept
Win Rate and Expected Value
Date: 2026-08-17
Most tests don’t win, and a programme’s value is the arithmetic across all of them rather than the story of the good ones. Getting this right reframes a run of failures from an embarrassment into the expected cost of the winners — which is the argument that keeps programmes funded.
Win rate is the share of completed tests that produce a significant positive result; expected value is the average gain per test once losses, nulls and wins are all counted.
The distribution
Reported win rates across mature programmes cluster somewhere around one in five to one in three, and lower on well-optimised surfaces. The distribution of effects is heavily skewed: many small nulls, a few modest wins, occasional large ones, and a meaningful number of losses you’re glad you caught.
100 tests run
22 won avg +1.8% ← ship
14 lost avg −2.4% ← don't ship. VALUE CREATED, see below
64 inconclusive ← the majority. not failures
value of the winners 22 × 1.8% = +39.6 percentage-points of relative gain
value of the losers 14 × 2.4% = +33.6 pp of loss AVOIDED
The losers are worth nearly as much as the winners, and this is the part that never makes it into a quarterly deck. Fourteen changes that would have been shipped on intuition, each costing ~2.4%, were stopped. That’s the programme’s second product and it’s invisible unless someone counts it — Inconclusive Results.
The expected value calculation
Per test, before running it:
page annual revenue £2,400,000
plausible relative lift if it works +2%
probability it works (your own win rate) 25%
expected value = 2,400,000 × 0.02 × 0.25 = £12,000/yr
cost = 6 dev-days + 3 weeks traffic ≈ £4,000
─────────
3× return
The honest input is the 25% — use your programme’s measured win rate, not optimism. This is the strongest practical argument for maintaining an Experiment Archive: without it, the base rate is a guess, and the guess is always too high.
At programme level:
48 tests/year × £12,000 expected value each = £576,000
programme cost (2 FTE + tooling) ≈ £220,000
─────────
2.6× return
This is the number to defend a budget with, and it survives a bad quarter in a way “we got a 6% win on the PDP” doesn’t.
Why the wins underdeliver
Two effects, both systematic, and both need pricing into the numbers above.
Winner’s Curse. You selected tests for having crossed a significance line, and crossing a line is easier with a favourable roll. The measured effect of a shipped winner is biased upward — severely so in underpowered tests, mildly in well-powered ones.
test power typical inflation of the measured effect
80% ~10–15% overstated
50% ~30–40% overstated
20% can be 2–3× overstated
Discount shipped wins before adding them to a total. A programme that sums raw test results will claim a cumulative uplift the site’s actual revenue doesn’t show, and that discrepancy is the fastest way to lose credibility.
Effects decay. Novelty fades, seasons change, and the site around the change moves on. A win measured in March is not necessarily still worth that in September — Novelty and Primacy Effects, Post-Test Validation.
The rigorous correction for both is a long-running holdout: a slice of users who never receive any shipped winner, compared against everyone else after a year. It’s the only measurement that catches inflation and decay together, and the gap between “sum of test results” and “holdout difference” is routinely large.
Why chasing win rate is a mistake
Win rate is trivially improvable by testing only safe things, and doing so destroys value.
CONSERVATIVE AMBITIOUS
40% win rate 18% win rate
avg win +0.6% avg win +4.2%
30 tests → 12 wins × 0.6% = 7.2 30 tests → 5.4 wins × 4.2% = 22.7
better win rate, third the value
Test bold changes. The distribution is skewed, so the programme’s return is dominated by its largest wins — and you can’t get a large win from a small change. A run of nulls on ambitious hypotheses is a healthier programme than a run of marginal wins on button colours.
Corollary: a programme reporting a 70% win rate has a measurement problem. The realistic explanations are peeking, no correction for multiplicity, post-hoc segmentation, or counting inconclusive results as wins — Peeking, Segmentation (test results).
What to report
- All tests, including the abandoned ones. Selective reporting is what makes the numbers untrustworthy
- Losses avoided, in pounds. The invisible half of the value
- Win rate as context, never as a target
- Shipped-win estimates discounted for winner’s curse, with the discount stated
- Holdout-measured cumulative effect annually, as the honest total — and expect it to be smaller than the sum of the parts
- Learning that transferred, which doesn’t fit an arithmetic but is often the largest output — Institutional Learning
Where it interacts
- Experimentation Velocity — the number of attempts, which multiplies everything here
- Test Prioritisation — expected value is the only scoring input with real numbers in it
- Winner’s Curse — the mechanism behind the discount, and why power is the lever that fixes it
- Holdout Groups — the measurement that makes the cumulative claim defensible