Tags: statistics concept
Confidence Intervals
Date: 2026-08-16
The range of effects your data is consistent with. It carries everything a p-value carries plus the two things it doesn’t — how big, and how precisely you know it — which is why it should be the number you report.
What it is
A confidence interval is a range built so that, across many repeated experiments, a stated proportion of the intervals produced would contain the true value. At 95%, nineteen intervals in twenty would capture it.
Worked, on a test that came back positive. Control 3.00%, variant 3.09%, 53,000 per arm:
difference = 0.09 pp
SE of difference = √[2 × 0.03045 × 0.96955 ÷ 53,000] = 0.1056 pp
95% CI = 0.09 ± 1.96 × 0.1056
= 0.09 ± 0.207
= −0.117 pp to +0.297 pp
As a relative lift on a 3% baseline: −3.9% to +9.9%.
The interval spans zero, so the result is not significant — but the interval says far more than that. It says the variant might be nearly 4% worse or nearly 10% better, and the data cannot distinguish between those. “No significant difference” collapses all of that into two words.
What 95% does not mean
In plain terms: the 95% describes the method, not this particular interval. It says the procedure produces intervals that capture the truth 19 times in 20. It does not say there’s a 95% chance the true value is inside the one in front of you — under the frequentist frame the true value is fixed and either is or isn’t in there.
The interval that does mean what people want is a credible interval, and it needs a prior — see Credible Intervals and Bayesian vs Frequentist. In practice most people read a confidence interval as a credible interval and are usually not badly misled, but the distinction matters when the stakes or the priors are unusual.
Reading one
Three things at a glance:
| What you see | What it means |
|---|---|
| Interval excludes zero | Significant at the matching level. Same information as p < α |
| Interval width | Precision. Wide means you learned little, whatever the point estimate says |
| Where the bounds sit commercially | The actual decision input |
The bounds are the decision, not the midpoint. From Guide - Statistics for CRO: a +3.0% point estimate with a +0.3% to +5.7% interval, on 500,000 annual visitors at £50, is a range of £2,250 to £42,750 a year. Against a £15,000 build, the lower bound loses money and the upper bound is excellent. That’s the finding — not “we made £22,500”.
Why it beats the p-value
- It separates size from certainty. A p-value fuses them, so a tiny effect on huge traffic and a large effect on modest traffic can produce the same number
- It makes null results informative. −0.4% to +0.6% has ruled out anything meaningful. −8% to +9% has ruled out nothing. Both report as “not significant”
- It translates to money directly. Multiply both bounds by traffic and margin
- It resists the dichotomy problem. No cliff at 0.05 — just a range that gets narrower with more data
Where it misleads
- Overlapping intervals do not mean no difference. Two intervals can overlap while the difference between them is significant. Test the difference; don’t eyeball two separate bars
- It assumes the analysis was planned. After Peeking or post-hoc segmentation, the coverage guarantee is gone — the interval is drawn with the same formula and no longer covers what it claims
- Ratio metrics need care. The variance of a ratio isn’t the variance of its parts, and naive intervals on them are too narrow — Ratio Metrics, Bootstrapping
- Heavy tails break the normal approximation. For revenue, Bootstrapping gives an honest interval where the formula gives a confident wrong one — Skewed and Heavy-Tailed Distributions