Tags: statistics concept
P-Values
Date: 2026-08-16
The probability of seeing data at least this extreme, if the null hypothesis were true. That conditional clause is the entire concept, and dropping it produces every misreading there is.
What it is
A p-value is one number answering one question: assuming there is genuinely no difference, how often would random sampling hand me a gap this large or larger?
Small p-value → the data would be surprising under the null → the null looks like a poor explanation.
Worked, using Hypothesis Testing’s example. Control 3.00%, variant 3.09%, 53,000 per arm, standard error of the difference 0.1056 percentage points:
observed difference = 0.09 pp
z = 0.09 ÷ 0.1056 = 0.85
p (two-tailed) = 0.395
p = 0.40. If the two versions were identical, you’d see a gap this big or bigger about 40% of the time. Entirely ordinary.
Now the same difference on 500,000 per arm — ten times the traffic:
SE = √[2 × 0.03045 × 0.96955 ÷ 500,000] = 0.000344 → 0.0344 pp
z = 0.09 ÷ 0.0344 = 2.62
p = 0.009
Identical effect, p = 0.009. The p-value moved from “nothing here” to “highly significant” without the effect changing at all. A p-value is a statement about evidence, not about size — it conflates how big the effect is with how much data you collected, and you cannot recover either from it alone. That’s the argument for reading Confidence Intervals and Effect Size instead.
The four misreadings
Each of these is wrong, and each is said constantly.
| The claim | Why it’s wrong |
|---|---|
| ”p = 0.03, so there’s a 3% chance the null is true” | The p-value is computed assuming the null is true. It can’t also be the probability of it. Getting that number needs a prior — Bayesian vs Frequentist |
| ”p = 0.03, so there’s a 97% chance the variant is better” | Same inversion. See Base Rate Fallacy |
| ”p = 0.001 means a bigger effect than p = 0.04” | It means more evidence, which is mostly a statement about sample size |
| ”p = 0.03 means it’d replicate 97% of the time” | Replication probability is much lower than intuition suggests, and isn’t what’s being computed |
In plain terms: the p-value tells you how weird your data would be in a world where your change did nothing. It does not tell you the chance you’re in that world, and it does not tell you how much your change is worth.
Why the intuition fails
The mind wants P(hypothesis | data). The p-value is P(data | hypothesis). Swapping the two feels harmless and isn’t.
The standard illustration: a disease affecting 1 in 1,000, a test that catches every case but fires falsely 5% of the time. You test positive.
per 100,000 people
100 have it → 100 test positive (true)
99,900 don't → 4,995 test positive (false)
─────
5,095 positives, of which 4,995 are wrong
The probability the test fires given you’re healthy is 5%. The probability you’re healthy given a positive test is 98%. Same two numbers, opposite conclusions, and only the base rate separates them.
The same structure applies to tests: if most of your ideas don’t work — and most don’t — then a meaningful share of your significant results are false positives regardless of your threshold.
Practical consequences
- Never report a p-value alone. Report the effect and its interval; the p-value adds little that the interval doesn’t carry better
- Don’t rank findings by p-value. p = 0.001 on a trivial effect beats nothing
- A p-value is only valid for the analysis you planned. Computed after choosing metrics, segments or a stopping point on the basis of results, it describes nothing at all — see Peeking, The Multiple Comparisons Problem and The Garden of Forking Paths
- 0.05 is a convention, not a discovery. See Statistical Significance