Tags: statistics experimentation concept
Practical vs Statistical Significance
Date: 2026-08-16
Statistical significance says the effect probably isn’t zero. Practical significance asks whether it’s big enough to be worth having. They come apart in both directions, and a programme can pass every statistical check for a year while producing nothing.
The two questions
- Statistically significant — the data would be unlikely if there were no effect. A claim about evidence
- Practically significant — the effect is large enough to change a decision. A claim about money
Nothing in the test computes the second. It’s your job, and it requires a number you set before the test: the smallest effect worth having. That number is also your MDE, which is why choosing it well does double duty.
Direction one: significant but worthless
With enough traffic, any non-zero difference crosses the line. On 5,000,000 visitors per arm:
control 3.000%
variant 3.012% +0.4% relative
SE = √[2 × 0.03006 × 0.96994 ÷ 5,000,000] = 0.01079 pp
z = 0.012 ÷ 0.01079 = 1.11
Not quite there. Push to 20,000,000 per arm and the standard error halves again, z reaches 2.22, and the same 0.4% relative lift lands at p ≈ 0.03 — significant, and worth about £3,000 a year on 500,000 annual visitors at £50.
In plain terms: significance measures how sure you are that the effect isn’t exactly zero. On very large samples you become very sure about effects too small to care about. Being certain about something trivial is still trivial.
Direction two: worthwhile but not significant
Far more common, and far more costly, because it looks like a negative result.
From the running example — 53,000 per arm, +3.0% relative, CI −3.9% to +9.9%. Not significant. But on 500,000 annual visitors at £50:
| Relative lift | Extra orders/year | Value | |
|---|---|---|---|
| CI lower | −3.9% | −585 | −£29,250 |
| Point estimate | +3.0% | +450 | +£22,500 |
| CI upper | +9.9% | +1,485 | +£74,250 |
The honest summary is “somewhere between losing £29,000 and gaining £74,000 a year” — which is not “no effect”, it’s “we have learned almost nothing and need more traffic”. Reporting it as a failed test discards a possibly valuable change and, worse, records a wrong belief in the archive.
See Inconclusive Results and Statistical Power.
The decision rule
Set the threshold first, then compare the interval to it — not to zero.
worth having = £15,000/year (build cost, paid back in year one)
≈ 300 extra orders
≈ 2% relative lift
Then read the interval against 2%, not against 0%:
CI entirely above 2% → ship
CI entirely below 2% → don't, even if significant
CI straddles 2% → more traffic, or a judgement call
This is the single change that most improves how test results get used, and it costs nothing but deciding the number in advance.
Where it bites
- Large-traffic sites over-ship. Everything is significant, so significance stops discriminating and the threshold has to become commercial
- Small-traffic sites under-ship. Nothing is significant, so real wins get binned. See Holdout Groups for measuring the cumulative effect of changes too small to test individually
- Guardrails invert it. For harm you care about small effects being ruled out, not detected. A guardrail with a wide interval hasn’t cleared anything — Guardrail Metrics
- Stakeholders hear “significant” as “important”. The word does the damage. Report effects and pounds; use the word sparingly — Communicating Uncertainty