Tags: statistics experimentation concept

Practical vs Statistical Significance

Date: 2026-08-16


Statistical significance says the effect probably isn’t zero. Practical significance asks whether it’s big enough to be worth having. They come apart in both directions, and a programme can pass every statistical check for a year while producing nothing.


The two questions

  • Statistically significant — the data would be unlikely if there were no effect. A claim about evidence
  • Practically significant — the effect is large enough to change a decision. A claim about money

Nothing in the test computes the second. It’s your job, and it requires a number you set before the test: the smallest effect worth having. That number is also your MDE, which is why choosing it well does double duty.

Direction one: significant but worthless

With enough traffic, any non-zero difference crosses the line. On 5,000,000 visitors per arm:

control  3.000%
variant  3.012%          +0.4% relative

SE = √[2 × 0.03006 × 0.96994 ÷ 5,000,000]  = 0.01079 pp
z  = 0.012 ÷ 0.01079                       = 1.11

Not quite there. Push to 20,000,000 per arm and the standard error halves again, z reaches 2.22, and the same 0.4% relative lift lands at p ≈ 0.03 — significant, and worth about £3,000 a year on 500,000 annual visitors at £50.

In plain terms: significance measures how sure you are that the effect isn’t exactly zero. On very large samples you become very sure about effects too small to care about. Being certain about something trivial is still trivial.

Direction two: worthwhile but not significant

Far more common, and far more costly, because it looks like a negative result.

From the running example — 53,000 per arm, +3.0% relative, CI −3.9% to +9.9%. Not significant. But on 500,000 annual visitors at £50:

Relative liftExtra orders/yearValue
CI lower−3.9%−585−£29,250
Point estimate+3.0%+450+£22,500
CI upper+9.9%+1,485+£74,250

The honest summary is “somewhere between losing £29,000 and gaining £74,000 a year” — which is not “no effect”, it’s “we have learned almost nothing and need more traffic”. Reporting it as a failed test discards a possibly valuable change and, worse, records a wrong belief in the archive.

See Inconclusive Results and Statistical Power.

The decision rule

Set the threshold first, then compare the interval to it — not to zero.

worth having  = £15,000/year   (build cost, paid back in year one)
             ≈ 300 extra orders
             ≈ 2% relative lift

Then read the interval against 2%, not against 0%:

  CI entirely above 2%     →  ship
  CI entirely below 2%     →  don't, even if significant
  CI straddles 2%          →  more traffic, or a judgement call

This is the single change that most improves how test results get used, and it costs nothing but deciding the number in advance.

Where it bites

  • Large-traffic sites over-ship. Everything is significant, so significance stops discriminating and the threshold has to become commercial
  • Small-traffic sites under-ship. Nothing is significant, so real wins get binned. See Holdout Groups for measuring the cumulative effect of changes too small to test individually
  • Guardrails invert it. For harm you care about small effects being ruled out, not detected. A guardrail with a wide interval hasn’t cleared anything — Guardrail Metrics
  • Stakeholders hear “significant” as “important”. The word does the damage. Report effects and pounds; use the word sparingly — Communicating Uncertainty