Tags: statistics experimentation concept
The Multiple Comparisons Problem
Date: 2026-08-16
Test twenty things at 95% confidence and you should expect one false positive. Not as a risk — as the specification. The threshold controls the error rate per comparison, and nobody makes only one comparison.
What it is
The multiple comparisons problem is that your significance threshold governs each comparison separately, so the chance of at least one false positive across a set of comparisons climbs with how many you make.
Set α = 0.05 and you’ve accepted a 1-in-20 false positive rate per test. Make twenty tests and the rate for the set as a whole is nothing like 1 in 20.
The arithmetic
Each independent test at α = 0.05 has a 95% chance of not producing a false positive. Across k tests:
| Comparisons | Chance of ≥1 false positive |
|---|---|
| 1 | 5% |
| 5 | 23% |
| 10 | 40% |
| 20 | 64% |
| 50 | 92% |
Working for k = 20:
P(none) = 0.95²⁰ = 0.358
P(at least 1) = 1 − 0.358 = 64.2%
In plain terms: if you look at twenty things that are all doing nothing, it’s more likely than not that at least one of them will look like it’s doing something. Finding a winner among twenty comparisons is the expected outcome of pure noise, not evidence against it.
The rate that matters here has a name: the family-wise error rate — the probability of at least one false positive across the whole set of comparisons, as opposed to the per-comparison α you set.
Where the comparisons hide
The count is almost always higher than people think, because most comparisons aren’t labelled as tests.
- Metrics. One test reporting conversion, AOV, revenue per visitor, bounce, add-to-cart and pages per session is six comparisons
- Segments. Mobile, desktop, tablet, new, returning, paid, organic — post-hoc slicing multiplies fast, and slicing two ways multiplies again
- Variants. A four-way test is three comparisons against control, not one
- Time. Checking each week and reporting the week it won is the same problem in a trench coat — see Peeking
- Across the programme. Fifty tests a year at 5% is two or three false winners in the archive, forever, unless replicated
The unlabelled ones are the dangerous ones. Nobody thinks “I have made twenty-eight comparisons” while clicking through a results dashboard.
Corrections
Two families, and the choice is about what you’re protecting against.
| Controls | Effect | Use when | |
|---|---|---|---|
| Bonferroni | Family-wise error rate | Divide α by k. Very conservative — costs a lot of power | Few comparisons, false positives expensive |
| Benjamini–Hochberg | False discovery rate — the proportion of your positives that are false | Ranks p-values and applies a sliding threshold. Much less power loss | Many comparisons, exploratory work |
Bonferroni worked, for six metrics:
α_adjusted = 0.05 ÷ 6 = 0.0083
Every metric now needs p < 0.0083. The critical value rises from 1.96 to 2.64, so the sample size scales by (2.64 + 0.84)² ÷ (1.96 + 0.84)² = 1.5× — 53,000 per arm becomes about 82,000. That’s the honest price of looking at six things, and it’s usually the argument for not looking at six things.
See Bonferroni and False Discovery Rate.
The better fix
Correction is the answer when multiple comparisons are unavoidable. Mostly they’re avoidable, and the design fix is cheaper than the statistical one:
- One primary metric, nominated before launch. Everything else is diagnostic and explicitly not decision-eligible — Overall Evaluation Criterion
- Pre-register segments. A segment named in the plan is a hypothesis. A segment found afterwards is a fishing expedition — Pre-Registration
- Treat surprises as hypotheses, not findings. A post-hoc pattern is a candidate for the next test, never a conclusion from this one
- Replicate anything surprising. Replication is the only reliable filter, and it’s cheaper than a corrected design
Ignoring all of this and reporting whatever won is P-Hacking when deliberate and The Garden of Forking Paths when not — and the second is far more common among people acting in good faith.