Tags: statistics concept

Statistical Significance

Date: 2026-08-16


A threshold decision, not a discovery. “Significant” means the p-value fell below a line someone chose by convention in the 1920s — it carries no information about whether the effect matters.


What it is

Statistically significant means the p-value came in below a threshold you chose before running the test. That is the whole content of the term — it is a comparison against a line, and it says nothing about size, importance or replicability.

The mechanism

Before the test, pick a significance level — written α, conventionally 0.05. It is the false positive rate you’re willing to run at: the proportion of tests where you’d wrongly reject a true null.

After the test, compare: p < α → “statistically significant” → reject the null.

A comparison against a number you chose in advance, and nothing more.

Where 0.05 came from

Fisher suggested it as a rough convenience in the 1920s, explicitly as a matter of taste rather than a finding, and it stuck. There is no mathematical or empirical basis for 0.05 being the right line, and the fields that use different lines aren’t wrong — particle physics runs at roughly one in 3.5 million, because a false discovery there costs more.

What α actually encodes is the cost ratio between the two mistakes. Shipping something useless versus missing something valuable. If shipping is cheap and reversible — a copy change behind a feature flag — 0.05 is conservative. If it’s a six-month rebuild, it’s reckless.

The dichotomy problem

p = 0.049   →  "significant"      →  ship it, write it up
p = 0.051   →  "not significant"  →  bin it, move on

These two results are indistinguishable as evidence. The difference between them is far smaller than the difference between p = 0.049 and p = 0.001, yet the threshold puts the first pair on opposite sides of a decision and the second pair on the same side.

Nothing in the maths creates that cliff. It’s an artefact of compressing a continuous measure of evidence into a binary, and it’s why reporting the interval alongside is not optional.

In plain terms: significance is a yes/no answer to a question that has a sliding-scale answer. Two tests can be nearly identical in strength of evidence and land on opposite sides of the line.

Significant does not mean important

The word does the damage. In ordinary English “significant” means consequential; in statistics it means detectable.

With enough traffic, any non-zero difference becomes significant. A 0.001 percentage point lift is real, detectable at ten million visitors, and worth nothing. See Practical vs Statistical Significance.

Running the relationship the other way is just as common: an effect worth £200,000 a year that came back non-significant because the test was underpowered. The effect didn’t fail to exist — the test failed to see it. See Statistical Power and Inconclusive Results.

Using it honestly

  • Set α before the test, and don’t move it. Adjusting the threshold after seeing the p-value is P-Hacking
  • Adjust it when you test several things, or the family-wise error rate climbs — The Multiple Comparisons Problem
  • Report the p-value, the effect and the interval together. “Significant” alone is a claim stripped of everything decision-relevant
  • Treat it as one input. Significance says the effect is probably not zero. Whether to ship depends on size, cost, risk and confidence — and only the first of those is in the test