Tags: statistics concept
Bonferroni and False Discovery Rate
Date: 2026-08-17
The two families of correction for testing many things at once. Bonferroni controls the chance of any false positive and is brutal; false discovery rate controls the proportion of your findings that are false and is usually the right trade for commercial work.
Bonferroni correction divides the significance threshold by the number of comparisons; false discovery rate (FDR) control instead caps the expected share of significant results that are false.
What each controls
The distinction is the whole note, and it’s a choice about which kind of mistake you’re protecting against.
BONFERRONI → family-wise error rate (FWER)
P(at least one false positive among all tests) ≤ α
"I want to be 95% sure NOTHING here is spurious"
BENJAMINI-HOCHBERG → false discovery rate (FDR)
E[proportion of my rejections that are false] ≤ α
"I accept that ~5% of what I flag will be wrong"
In plain terms: Bonferroni protects you from ever crying wolf. FDR accepts that one in twenty of your alarms is false, in exchange for hearing far more of the real ones. Which you want depends on what a false positive costs — The Multiple Comparisons Problem.
Bonferroni, worked
Divide α by the number of tests.
20 metrics, α = 0.05
adjusted threshold = 0.05 / 20 = 0.0025
p-value vs 0.05 vs 0.0025
0.001 significant significant ✓ survives
0.008 significant not ✗
0.012 significant not ✗
0.030 significant not ✗
Three findings that looked significant are discarded. The family-wise error rate is now controlled:
P(no false positive) = (1 − 0.0025)²⁰ = 0.951 → FWER ≈ 4.9% ≤ 5% ✓
The cost is power, and it’s severe. Detecting the same effect at α = 0.0025 instead of 0.05 needs roughly 1.7× the sample size per test. With 20 metrics you’d need nearly double the traffic to keep the same ability to detect a real effect — which is why Bonferroni applied to a dashboard of metrics means detecting almost nothing.
Holm-Bonferroni is a strictly better version — same guarantee, more power — and there’s no reason to use plain Bonferroni over it except that plain Bonferroni is easier to explain. Sort p-values ascending, compare the ith against α / (m − i + 1), stop at the first failure.
Benjamini-Hochberg, worked
Sort ascending, compare each p-value against a threshold that rises with its rank.
m = 10 tests, α = 0.05
threshold for rank i = (i / m) × α
rank i p-value (i/10) × 0.05 p ≤ threshold?
1 0.001 0.005 yes
2 0.008 0.010 yes
3 0.012 0.015 yes
4 0.021 0.020 no
5 0.030 0.025 no
6 0.041 0.030 no
7 0.055 0.035 no
8 0.210 0.040 no
9 0.400 0.045 no
10 0.780 0.050 no
find the LARGEST rank where the answer is yes → rank 3
reject every hypothesis up to and including rank 3 → 3 discoveries
The step people get wrong: you don’t reject only the ones that individually pass. You find the largest passing rank and reject everything at or below it. If rank 4 had passed while rank 3 failed, you would still reject ranks 1 through 4.
Compare the two on the same data:
Bonferroni 0.05 / 10 = 0.005 → only p = 0.001 survives → 1 discovery
Benjamini-Hochberg → ranks 1–3 survive → 3 discoveries
BH found 3× as much, at the cost of expecting ~5% of its
findings (so, ~0.15 of these 3) to be false
Choosing
| Use Bonferroni / Holm | Use BH / FDR | |
|---|---|---|
| A false positive costs | A great deal — a shipped harmful change, a regulatory claim | Some wasted follow-up |
| Number of tests | Few (2–10) | Many (10+) |
| Purpose | Confirmatory — deciding | Exploratory — generating leads |
| In practice here | Guardrail Metrics — no, wait, see below | Diagnostic metrics, segment scans, multi-arm tests |
Guardrails are the interesting exception. For guardrails you deliberately don’t want to correct, or you want to correct in the opposite direction — the expensive error there is missing real harm, not raising a false alarm. Correcting guardrails makes them less sensitive to exactly what they exist to catch. Leave them uncorrected and accept the false alarms.
What doesn’t need correcting
Over-applying corrections is its own failure, and produces the “we can’t detect anything” outcome:
- One pre-registered primary metric. n = 1, nothing to correct — which is precisely why the one-primary-metric rule exists. It’s a multiplicity control disguised as a discipline
- Diagnostics that can’t change the decision. If nothing is being rejected, there’s no error rate to control — Secondary and Diagnostic Metrics
- Guardrails, as above
- Independent tests answering unrelated questions. Correcting across every test your company runs this year is not required — the family is the set of comparisons bearing on one decision
Deciding the family is the judgement call, and it’s where the honesty lives. If you’d have shipped on whichever of four metrics won, those four are one family.
The failure this doesn’t fix
Corrections handle the comparisons you report. They do nothing about the ones you ran and discarded, or the analytical choices you made along the way — which is a much larger and less visible problem — The Garden of Forking Paths, P-Hacking.
Pre-registration is what makes a correction meaningful, because it fixes m before you look — Pre-Registration.
Where it interacts
- The Multiple Comparisons Problem — the arithmetic these correct for
- Statistical Power — what corrections cost, and why the sample size has to be planned with m in mind
- Segmentation (test results) — the commonest place corrections are needed and skipped
- Peeking — multiplicity over time rather than across metrics, needing a different fix — Sequential Testing