Tags: statistics concept

Base Rate Fallacy

Date: 2026-08-17


Reading a test’s accuracy as the probability it’s right, while ignoring how rare the thing is. When something is uncommon, even a very accurate detector produces mostly false alarms — and the intuition that a “95% accurate” test is 95% likely to be correct is wrong by an order of magnitude.


The base rate fallacy is judging how likely something is from a test’s accuracy alone, ignoring the base rate — how common the thing was before you tested.

The worked case

A fraud detection model, described as 95% accurate in both directions.

sensitivity  95%   catches 95% of genuine fraud
specificity  95%   correctly clears 95% of legitimate orders
base rate    0.5%  of orders are fraudulent

Take 10,000 orders and count:

                        actually fraud    actually legitimate    total
                              50                 9,950         10,000

flagged by the model          47.5              497.5           545
  (50 × 0.95)                                (9,950 × 0.05)

not flagged                    2.5            9,452.5         9,455
P(fraud | flagged)  =  47.5 / 545  =  0.087   →  8.7%

In plain terms: the model is right about 95% of the orders it sees, and when it raises an alarm there’s only about a 9% chance the order is actually fraudulent. Over 91% of its alerts are false — not because the model is bad, but because legitimate orders outnumber fraudulent ones 199 to 1, so even a 5% error rate on that huge group swamps the small group entirely.

The 497.5 is the whole lesson. Five percent of a very large number is much bigger than ninety-five percent of a very small one.

Why the intuition fails

People substitute one conditional probability for its reverse:

what you're TOLD          P(flagged | fraud)      =  95%
what you WANT             P(fraud | flagged)      =  8.7%

these are different quantities and are not interchangeable

The link between them is Bayes’ theorem, and the base rate is the term intuition drops:

P(A|B)  =  P(B|A) × P(A) / P(B)

P(fraud | flagged)  =  0.95 × 0.005 / 0.0545
                    =  0.00475 / 0.0545
                    =  0.087    ✓ matches the count above

Counting in a table beats the formula for most purposes — it’s harder to get wrong and much easier to explain to someone else.

How much the base rate moves it

Same model, different prevalence:

base rate   flagged, true    flagged, false    P(fraud | flagged)
  0.1%           9.5              999.0             0.9%
  0.5%          47.5              497.5             8.7%
  2.0%         190.0              490.0            27.9%
 10.0%         950.0              450.0            67.9%
 30.0%       2,850.0              350.0            89.1%

Nothing about the model changed across those rows. The same detector is useless at 0.1% prevalence and good at 30%. This is why a model’s reported accuracy is close to meaningless without the base rate of the population it will run on — and why a model validated on a balanced dataset behaves completely differently in production.

Where it bites in this work

  • Fraud and bot detection. Both target rare events, so both generate mostly false positives. A review queue built on the arithmetic above needs staffing for 545 reviews to catch 47 frauds — Bot and Internal Traffic
  • Anomaly alerting. Watching 40 metrics hourly at a 1% false positive rate produces about 10 false alarms a day, because genuine breaks are rare. This is the arithmetic behind alert fatigue — Anomaly Detection, Alerting on Metrics
  • Reading a “significant” test result. The most consequential application:
if only 20% of your test ideas genuinely work

  100 tests:  20 real effects, 80 null

  at 80% power and α = 0.05
    true positives   20 × 0.80  =  16
    false positives  80 × 0.05  =   4

  P(real | significant)  =  16 / 20  =  80%

now the same programme at 30% power (underpowered tests)
    true positives   20 × 0.30  =   6
    false positives  80 × 0.05  =   4

  P(real | significant)  =   6 / 10  =  60%

Four in ten of your “winners” are noise, in the underpowered programme. The p-value didn’t change — the proportion of your significant results that are real collapsed because power fell. This is the strongest available argument for adequate power, and it’s the reason a run of underpowered tests actively degrades decision quality rather than merely being inefficient — Statistical Power, Winner’s Curse.

Lowering your prior further makes it worse. If only 10% of ideas work, the same 80%-powered programme gives 8 / (8 + 4.5) = 64% — so a testing programme’s credibility depends as much on idea quality as on statistical rigour.

  • Churn and propensity models targeting rare outcomes, where a “90% accurate” churn model flags mostly people who won’t churn — Logistic Regression

The habit that prevents it

Whenever handed an accuracy figure, ask for the base rate and count in a table of 10,000. It takes a minute, needs no formula, and it’s the single most reliable defence against this — including against your own reasoning, since knowing about the fallacy provides very little protection from committing it.

Where it interacts

  • Prior Likelihood and Posterior — the base rate is a prior, and this is the same update in its starkest form
  • P-Values — the misreading this note quantifies: p is P(data | no effect), never P(no effect | data)
  • Statistical Power — the lever that determines what fraction of your significant results are real
  • Choosing a Prior — where a realistic base rate gets encoded deliberately rather than ignored