Tags: statistics concept

Bonferroni and False Discovery Rate

Date: 2026-08-17


The two families of correction for testing many things at once. Bonferroni controls the chance of any false positive and is brutal; false discovery rate controls the proportion of your findings that are false and is usually the right trade for commercial work.


Bonferroni correction divides the significance threshold by the number of comparisons; false discovery rate (FDR) control instead caps the expected share of significant results that are false.

What each controls

The distinction is the whole note, and it’s a choice about which kind of mistake you’re protecting against.

BONFERRONI  →  family-wise error rate (FWER)
               P(at least one false positive among all tests) ≤ α
               "I want to be 95% sure NOTHING here is spurious"

BENJAMINI-HOCHBERG  →  false discovery rate (FDR)
               E[proportion of my rejections that are false] ≤ α
               "I accept that ~5% of what I flag will be wrong"

In plain terms: Bonferroni protects you from ever crying wolf. FDR accepts that one in twenty of your alarms is false, in exchange for hearing far more of the real ones. Which you want depends on what a false positive costs — The Multiple Comparisons Problem.

Bonferroni, worked

Divide α by the number of tests.

20 metrics, α = 0.05

adjusted threshold  =  0.05 / 20  =  0.0025

p-value    vs 0.05      vs 0.0025
0.001      significant  significant       ✓ survives
0.008      significant  not               ✗
0.012      significant  not               ✗
0.030      significant  not               ✗

Three findings that looked significant are discarded. The family-wise error rate is now controlled:

P(no false positive) = (1 − 0.0025)²⁰ = 0.951  →  FWER ≈ 4.9% ≤ 5%  ✓

The cost is power, and it’s severe. Detecting the same effect at α = 0.0025 instead of 0.05 needs roughly 1.7× the sample size per test. With 20 metrics you’d need nearly double the traffic to keep the same ability to detect a real effect — which is why Bonferroni applied to a dashboard of metrics means detecting almost nothing.

Holm-Bonferroni is a strictly better version — same guarantee, more power — and there’s no reason to use plain Bonferroni over it except that plain Bonferroni is easier to explain. Sort p-values ascending, compare the ith against α / (m − i + 1), stop at the first failure.

Benjamini-Hochberg, worked

Sort ascending, compare each p-value against a threshold that rises with its rank.

m = 10 tests, α = 0.05
threshold for rank i  =  (i / m) × α

rank i   p-value    (i/10) × 0.05    p ≤ threshold?
  1       0.001        0.005              yes
  2       0.008        0.010              yes
  3       0.012        0.015              yes
  4       0.021        0.020              no
  5       0.030        0.025              no
  6       0.041        0.030              no
  7       0.055        0.035              no
  8       0.210        0.040              no
  9       0.400        0.045              no
 10       0.780        0.050              no

find the LARGEST rank where the answer is yes  →  rank 3
reject every hypothesis up to and including rank 3  →  3 discoveries

The step people get wrong: you don’t reject only the ones that individually pass. You find the largest passing rank and reject everything at or below it. If rank 4 had passed while rank 3 failed, you would still reject ranks 1 through 4.

Compare the two on the same data:

Bonferroni  0.05 / 10 = 0.005    →  only p = 0.001 survives   →  1 discovery
Benjamini-Hochberg                →  ranks 1–3 survive         →  3 discoveries

BH found 3× as much, at the cost of expecting ~5% of its
findings (so, ~0.15 of these 3) to be false

Choosing

Use Bonferroni / HolmUse BH / FDR
A false positive costsA great deal — a shipped harmful change, a regulatory claimSome wasted follow-up
Number of testsFew (2–10)Many (10+)
PurposeConfirmatory — decidingExploratory — generating leads
In practice hereGuardrail Metrics — no, wait, see belowDiagnostic metrics, segment scans, multi-arm tests

Guardrails are the interesting exception. For guardrails you deliberately don’t want to correct, or you want to correct in the opposite direction — the expensive error there is missing real harm, not raising a false alarm. Correcting guardrails makes them less sensitive to exactly what they exist to catch. Leave them uncorrected and accept the false alarms.

What doesn’t need correcting

Over-applying corrections is its own failure, and produces the “we can’t detect anything” outcome:

  • One pre-registered primary metric. n = 1, nothing to correct — which is precisely why the one-primary-metric rule exists. It’s a multiplicity control disguised as a discipline
  • Diagnostics that can’t change the decision. If nothing is being rejected, there’s no error rate to control — Secondary and Diagnostic Metrics
  • Guardrails, as above
  • Independent tests answering unrelated questions. Correcting across every test your company runs this year is not required — the family is the set of comparisons bearing on one decision

Deciding the family is the judgement call, and it’s where the honesty lives. If you’d have shipped on whichever of four metrics won, those four are one family.

The failure this doesn’t fix

Corrections handle the comparisons you report. They do nothing about the ones you ran and discarded, or the analytical choices you made along the way — which is a much larger and less visible problem — The Garden of Forking Paths, P-Hacking.

Pre-registration is what makes a correction meaningful, because it fixes m before you look — Pre-Registration.

Where it interacts