Tags: statistics experimentation concept
Probability to Beat Control
Date: 2026-08-17
The headline number most Bayesian test tools report: the probability the variant’s true rate exceeds the control’s. It’s genuinely easier to read than a p-value and it throws away the thing you most need — it says how likely a win is, and nothing about how big.
What it is
Probability to beat control is the Bayesian posterior probability that the variant’s true rate is higher than control’s.
P(B > A) = probability that the variant's true rate exceeds control's,
given the data and the priors
Computed by simulation from the two posteriors:
draw 100,000 pairs, one from each posterior, count how often B > A
control Beta(301, 9,701) mean 3.009%
variant Beta(340, 9,662) mean 3.400%
→ B exceeded A in 96,200 of 100,000 draws
→ P(B > A) = 96.2%
In plain terms: given what you’ve seen, there’s about a 96% chance the variant is genuinely better than the control by some amount, however tiny.
That last clause is doing all the damage.
The trap
A high P(B>A) is compatible with an effect too small to be worth anything.
TEST 1 TEST 2
P(B > A) = 96% P(B > A) = 96%
95% credible int = +0.04pp to 95% credible int = +0.9pp to
+0.74pp +3.1pp
in relative terms = +1.3% to +25% in relative terms = +30% to +100%
annual value ≈ £4,000 annual value ≈ £180,000
to £74,000 to £600,000
same headline number. completely different decisions.
P(B>A) answers “is it better?” and never “by how much?” Two tests with identical 96% figures can differ by an order of magnitude in value, and the number that distinguishes them is the credible interval — which is usually one click further into the interface than the big percentage on the summary screen.
The asymmetry gets worse as sample size grows. A very large test on a trivially better variant will report P(B>A) approaching 100%, because you’ve measured a tiny real effect precisely. P(B>A) → 100% is a statement about certainty, not about magnitude — the same failure mode as reading a p-value as an effect size — Practical vs Statistical Significance, Effect Size.
Expected loss, which is the better number
The decision-theoretic companion, and the one worth promoting to the headline.
Expected loss is how much you’d expect to give up by choosing a variant, averaged over the posterior — counting only the scenarios where you’d be wrong.
for each of 100,000 simulated pairs:
if B < A, the loss is (A − B); otherwise the loss is 0
expected loss = average of those losses
TEST 1 expected loss of shipping B = 0.0009pp
TEST 2 expected loss of shipping B = 0.0008pp
Set a threshold in units you care about, and ship when expected loss falls below it — commonly phrased as “the caring threshold”. This is a genuinely better rule than “ship at 95%”, because it prices being wrong rather than merely counting how often you would be.
✗ "ship when P(B > A) > 95%"
→ ships trivial wins; ignores magnitude entirely
✓ "ship when expected loss < 0.01pp of conversion"
→ answers: if I'm wrong, how much does it cost me?
Reading one properly
- Never read P(B>A) alone. Always with the credible interval, which is the number that carries the magnitude — Credible Intervals
- It is not a p-value, and doesn’t map onto one.
P(B>A) = 96%is notp = 0.04. They’re different quantities from different frameworks and coincide only approximately, and only with weak priors at large samples - It depends on the prior. At small samples, a different prior moves it materially — and most tools don’t show you which prior they used — Choosing a Prior
- 95% is not a law. It’s a threshold inherited by analogy from frequentist convention, and it has no special standing here. Choose it from the cost of being wrong
- It doesn’t validate the test. A broken split produces a confident P(B>A) — check SRM, guardrails and duration first, always — Reading a Test Result, Sample Ratio Mismatch
The stopping problem
The specific misuse that tools encourage: watching P(B>A) climb and stopping when it crosses 95%.
The posterior is valid at any moment; a rule that stops at the first favourable moment is not. Checking daily and stopping on the first crossing selects for random upward wanders, and it inflates the rate at which you ship null changes — in exactly the way Peeking does for frequentist tests, even though no p-value is being corrupted.
The fix is the same in either framework: a stopping rule fixed before launch — a sample size, a duration, or a design built for continuous monitoring — Stopping Rules, Always-Valid Inference.
Where it interacts
- Credible Intervals — the number that supplies what this one omits, and the one to report alongside it
- Bayesian vs Frequentist — why this reads more naturally than a p-value, and what that ease conceals
- Winner’s Curse — shipping on a threshold selects for favourable noise here exactly as it does in frequentist testing
- Overall Evaluation Criterion — expected loss only means something if it’s computed on the one metric you decided in advance