Tags: statistics concept

Type I and Type II Errors

Date: 2026-08-16


Two ways to be wrong, and you trade one against the other. Tightening your threshold to avoid false positives directly buys you more false negatives — there is no setting that reduces both except more traffic.


A Type I error is a false positive — declaring an effect that isn’t there, at rate α (alpha). A Type II error is a false negative — missing an effect that is, at rate β (beta). Power is 1 − β.

The four outcomes

                          THE TRUTH
                  no effect        real effect
              ┌───────────────┬───────────────────┐
   reject H₀  │  TYPE I       │  correct          │
  "we found   │  false        │  detection        │
   something" │  positive  α  │           1 − β   │
              ├───────────────┼───────────────────┤
   fail to    │  correct      │  TYPE II          │
   reject H₀  │               │  false            │
  "nothing    │      1 − α    │  negative      β  │
   here"      │               │                   │
              └───────────────┴───────────────────┘
  • Type I error — you ship a change that does nothing. Rate is α, the significance level you set, conventionally 5%
  • Type II error — you bin a change that would have worked. Rate is β, conventionally 20%, which makes power (1 − β) 80%

In plain terms: a Type I error is believing something that isn’t there. A Type II error is missing something that is. The first is embarrassing and visible; the second is invisible, which is why almost everyone under-weights it.

The trade

The two rates move in opposite directions when you change the threshold, and only in opposite directions.

αType I rateEffect on Type II
0.101 in 10fewer misses
0.051 in 20conventional
0.011 in 100many more misses

Demanding more evidence before believing something means you’ll believe fewer true things too. The only lever that improves both at once is sample size — more data narrows the distribution under both hypotheses, so they overlap less. That’s why Sample Size Calculation takes α and β as inputs rather than producing them.

Worked: what 5% actually costs

Twenty tests in a year, none of which does anything:

20 tests × 0.05 = 1 expected false positive

One “winner” per twenty dead tests, guaranteed, from a threshold working exactly as designed. Not a bug — the specification. This is the seed of The Multiple Comparisons Problem, and it’s why an experiment archive full of unreplicated single wins should be read sceptically.

Now the other side, at conventional 80% power:

10 tests with a real effect × 0.20 = 2 expected misses

Two working ideas discarded per ten, and you will never find out which two. Type II errors leave no evidence behind — the test just says “nothing here” and everyone moves on.

Which one should you fear

Depends entirely on what the mistake costs, and CRO’s usual answer is the opposite of academia’s.

  • Cheap, reversible change — a false positive costs a deploy and some wasted attention. A false negative costs a real gain, permanently. Being strict is the more expensive error
  • Expensive or risky change — a rebuild, a pricing change, anything touching trust or compliance. A false positive costs months. Be strict
  • Guardrail metrics invert the logic entirely. You are looking for harm, so a missed harm is the expensive error — which is why guardrails are watched at looser thresholds than the primary metric. See Guardrail Metrics

Where they hide

  • Underpowered tests inflate Type II massively, and report themselves as “no difference” with no warning attached — Statistical Power
  • Peeking inflates Type I far past the α you set, without changing the number in the report
  • Multiple metrics inflate Type I in proportion to how many you look at
  • The Winner’s Curse is the shadow of Type I: because you only ship things that crossed the line, your shipped set is enriched with lucky results, and their effects shrink on contact with reality