Tags: experimentation statistics concept

Win Rate and Expected Value

Date: 2026-08-17


Most tests don’t win, and a programme’s value is the arithmetic across all of them rather than the story of the good ones. Getting this right reframes a run of failures from an embarrassment into the expected cost of the winners — which is the argument that keeps programmes funded.


Win rate is the share of completed tests that produce a significant positive result; expected value is the average gain per test once losses, nulls and wins are all counted.

The distribution

Reported win rates across mature programmes cluster somewhere around one in five to one in three, and lower on well-optimised surfaces. The distribution of effects is heavily skewed: many small nulls, a few modest wins, occasional large ones, and a meaningful number of losses you’re glad you caught.

100 tests run

  22   won            avg  +1.8%   ← ship
  14   lost           avg  −2.4%   ← don't ship. VALUE CREATED, see below
  64   inconclusive                ← the majority. not failures

value of the winners     22 × 1.8%  =  +39.6 percentage-points of relative gain
value of the losers      14 × 2.4%  =  +33.6 pp of loss AVOIDED

The losers are worth nearly as much as the winners, and this is the part that never makes it into a quarterly deck. Fourteen changes that would have been shipped on intuition, each costing ~2.4%, were stopped. That’s the programme’s second product and it’s invisible unless someone counts it — Inconclusive Results.

The expected value calculation

Per test, before running it:

page annual revenue                      £2,400,000
plausible relative lift if it works              +2%
probability it works (your own win rate)         25%

expected value  =  2,400,000 × 0.02 × 0.25       =  £12,000/yr
cost            =  6 dev-days + 3 weeks traffic  ≈  £4,000
                                                    ─────────
                                                    3× return

The honest input is the 25% — use your programme’s measured win rate, not optimism. This is the strongest practical argument for maintaining an Experiment Archive: without it, the base rate is a guess, and the guess is always too high.

At programme level:

48 tests/year × £12,000 expected value each     =  £576,000
programme cost (2 FTE + tooling)                 ≈ £220,000
                                                   ─────────
                                                   2.6× return

This is the number to defend a budget with, and it survives a bad quarter in a way “we got a 6% win on the PDP” doesn’t.

Why the wins underdeliver

Two effects, both systematic, and both need pricing into the numbers above.

Winner’s Curse. You selected tests for having crossed a significance line, and crossing a line is easier with a favourable roll. The measured effect of a shipped winner is biased upward — severely so in underpowered tests, mildly in well-powered ones.

test power    typical inflation of the measured effect

80%           ~10–15% overstated
50%           ~30–40% overstated
20%           can be 2–3× overstated

Discount shipped wins before adding them to a total. A programme that sums raw test results will claim a cumulative uplift the site’s actual revenue doesn’t show, and that discrepancy is the fastest way to lose credibility.

Effects decay. Novelty fades, seasons change, and the site around the change moves on. A win measured in March is not necessarily still worth that in September — Novelty and Primacy Effects, Post-Test Validation.

The rigorous correction for both is a long-running holdout: a slice of users who never receive any shipped winner, compared against everyone else after a year. It’s the only measurement that catches inflation and decay together, and the gap between “sum of test results” and “holdout difference” is routinely large.

Why chasing win rate is a mistake

Win rate is trivially improvable by testing only safe things, and doing so destroys value.

CONSERVATIVE                         AMBITIOUS

40% win rate                         18% win rate
avg win +0.6%                        avg win +4.2%
30 tests → 12 wins × 0.6% = 7.2      30 tests → 5.4 wins × 4.2% = 22.7

better win rate, third the value

Test bold changes. The distribution is skewed, so the programme’s return is dominated by its largest wins — and you can’t get a large win from a small change. A run of nulls on ambitious hypotheses is a healthier programme than a run of marginal wins on button colours.

Corollary: a programme reporting a 70% win rate has a measurement problem. The realistic explanations are peeking, no correction for multiplicity, post-hoc segmentation, or counting inconclusive results as wins — Peeking, Segmentation (test results).

What to report

  • All tests, including the abandoned ones. Selective reporting is what makes the numbers untrustworthy
  • Losses avoided, in pounds. The invisible half of the value
  • Win rate as context, never as a target
  • Shipped-win estimates discounted for winner’s curse, with the discount stated
  • Holdout-measured cumulative effect annually, as the honest total — and expect it to be smaller than the sum of the parts
  • Learning that transferred, which doesn’t fit an arithmetic but is often the largest output — Institutional Learning

Where it interacts