Tags: statistics concept
Base Rate Fallacy
Date: 2026-08-17
Reading a test’s accuracy as the probability it’s right, while ignoring how rare the thing is. When something is uncommon, even a very accurate detector produces mostly false alarms — and the intuition that a “95% accurate” test is 95% likely to be correct is wrong by an order of magnitude.
The base rate fallacy is judging how likely something is from a test’s accuracy alone, ignoring the base rate — how common the thing was before you tested.
The worked case
A fraud detection model, described as 95% accurate in both directions.
sensitivity 95% catches 95% of genuine fraud
specificity 95% correctly clears 95% of legitimate orders
base rate 0.5% of orders are fraudulent
Take 10,000 orders and count:
actually fraud actually legitimate total
50 9,950 10,000
flagged by the model 47.5 497.5 545
(50 × 0.95) (9,950 × 0.05)
not flagged 2.5 9,452.5 9,455
P(fraud | flagged) = 47.5 / 545 = 0.087 → 8.7%
In plain terms: the model is right about 95% of the orders it sees, and when it raises an alarm there’s only about a 9% chance the order is actually fraudulent. Over 91% of its alerts are false — not because the model is bad, but because legitimate orders outnumber fraudulent ones 199 to 1, so even a 5% error rate on that huge group swamps the small group entirely.
The 497.5 is the whole lesson. Five percent of a very large number is much bigger than ninety-five percent of a very small one.
Why the intuition fails
People substitute one conditional probability for its reverse:
what you're TOLD P(flagged | fraud) = 95%
what you WANT P(fraud | flagged) = 8.7%
these are different quantities and are not interchangeable
The link between them is Bayes’ theorem, and the base rate is the term intuition drops:
P(A|B) = P(B|A) × P(A) / P(B)
P(fraud | flagged) = 0.95 × 0.005 / 0.0545
= 0.00475 / 0.0545
= 0.087 ✓ matches the count above
Counting in a table beats the formula for most purposes — it’s harder to get wrong and much easier to explain to someone else.
How much the base rate moves it
Same model, different prevalence:
base rate flagged, true flagged, false P(fraud | flagged)
0.1% 9.5 999.0 0.9%
0.5% 47.5 497.5 8.7%
2.0% 190.0 490.0 27.9%
10.0% 950.0 450.0 67.9%
30.0% 2,850.0 350.0 89.1%
Nothing about the model changed across those rows. The same detector is useless at 0.1% prevalence and good at 30%. This is why a model’s reported accuracy is close to meaningless without the base rate of the population it will run on — and why a model validated on a balanced dataset behaves completely differently in production.
Where it bites in this work
- Fraud and bot detection. Both target rare events, so both generate mostly false positives. A review queue built on the arithmetic above needs staffing for 545 reviews to catch 47 frauds — Bot and Internal Traffic
- Anomaly alerting. Watching 40 metrics hourly at a 1% false positive rate produces about 10 false alarms a day, because genuine breaks are rare. This is the arithmetic behind alert fatigue — Anomaly Detection, Alerting on Metrics
- Reading a “significant” test result. The most consequential application:
if only 20% of your test ideas genuinely work
100 tests: 20 real effects, 80 null
at 80% power and α = 0.05
true positives 20 × 0.80 = 16
false positives 80 × 0.05 = 4
P(real | significant) = 16 / 20 = 80%
now the same programme at 30% power (underpowered tests)
true positives 20 × 0.30 = 6
false positives 80 × 0.05 = 4
P(real | significant) = 6 / 10 = 60%
Four in ten of your “winners” are noise, in the underpowered programme. The p-value didn’t change — the proportion of your significant results that are real collapsed because power fell. This is the strongest available argument for adequate power, and it’s the reason a run of underpowered tests actively degrades decision quality rather than merely being inefficient — Statistical Power, Winner’s Curse.
Lowering your prior further makes it worse. If only 10% of ideas work, the same 80%-powered programme gives 8 / (8 + 4.5) = 64% — so a testing programme’s credibility depends as much on idea quality as on statistical rigour.
- Churn and propensity models targeting rare outcomes, where a “90% accurate” churn model flags mostly people who won’t churn — Logistic Regression
The habit that prevents it
Whenever handed an accuracy figure, ask for the base rate and count in a table of 10,000. It takes a minute, needs no formula, and it’s the single most reliable defence against this — including against your own reasoning, since knowing about the fallacy provides very little protection from committing it.
Where it interacts
- Prior Likelihood and Posterior — the base rate is a prior, and this is the same update in its starkest form
- P-Values — the misreading this note quantifies: p is
P(data | no effect), neverP(no effect | data) - Statistical Power — the lever that determines what fraction of your significant results are real
- Choosing a Prior — where a realistic base rate gets encoded deliberately rather than ignored