Tags: statistics concept
Credible Intervals
Date: 2026-08-17
The interval that means what everyone wrongly thinks a confidence interval means: given the data and the prior, there’s a 95% probability the true value lies inside it. Same numbers as a confidence interval on a large sample, radically easier to state without saying something false.
A credible interval is a range that, given the data and the prior, contains the true value with a stated probability — the Bayesian counterpart of a confidence interval.
The definition people want
CREDIBLE INTERVAL (Bayesian) CONFIDENCE INTERVAL (frequentist)
"there is a 95% probability the "if I repeated this experiment many
true conversion rate is between times and built an interval each
2.8% and 3.2%" time, 95% of those intervals would
contain the true rate. about THIS
interval, I say nothing probabilistic"
a statement about the PARAMETER a statement about the PROCEDURE
Both sentences describe the same two numbers. The first is what people say when they report a confidence interval, and it’s wrong — the frequentist parameter is fixed, so it’s either in the interval or it isn’t, and there’s no probability about it. The Bayesian version treats the parameter as uncertain, so the probability statement is licensed — Confidence Intervals, Bayesian vs Frequentist.
Reading one off a posterior
The posterior is a distribution; the interval is the middle 95% of it.
posterior for the variant: Beta(451, 14,551)
mean = 451 / 15,002 = 0.03006
approximate sd of a Beta:
√( αβ / ((α+β)²(α+β+1)) )
= √( (451 × 14,551) / (15,002² × 15,003) )
= √( 6,562,501 / (225,060,004 × 15,003) )
= √( 6,562,501 / 3,376,575,240,012 )
= √0.000001944
= 0.001394
95% credible interval ≈ 0.03006 ± (1.96 × 0.001394)
= 0.02733 to 0.03279
→ 2.73% to 3.28%
In plain terms: given the prior and 15,000 observations, there’s a 95% probability the true conversion rate is between 2.73% and 3.28%.
Two conventions for choosing which 95%:
- Equal-tailed — cut 2.5% off each end. What almost every tool reports, and what’s computed above
- Highest density (HDI) — the narrowest interval containing 95%. Identical for symmetric posteriors; better for skewed ones, since an equal-tailed interval on a skewed posterior can exclude the most probable values
For conversion rates at any reasonable sample size the posterior is near-symmetric and the two coincide. For revenue metrics they can differ noticeably — Skewed and Heavy-Tailed Distributions.
The interval on the difference
What you actually want in a test isn’t each arm’s interval — it’s the interval on the difference, and this is where the Bayesian version is genuinely more convenient.
control Beta(301, 9,701)
variant Beta(340, 9,662)
draw 100,000 samples from each posterior, subtract pairwise,
read the percentiles of the difference
2.5th percentile +0.04pp
median +0.39pp
97.5th percentile +0.74pp
→ 95% credible interval on the absolute lift: +0.04pp to +0.74pp
→ in relative terms, roughly +1.3% to +24.6%
No formula needed. With frequentist methods, the interval on a difference of proportions needs a derivation, and the interval on a ratio of proportions needs another; here you simulate and read percentiles. The same trick handles any derived quantity — expected revenue impact, difference in ratios, anything — which is the practical reason Bayesian tooling is pleasant to work with.
Reading one honestly
- Width is the information.
+0.04pp to +0.74ppis a real but imprecisely-measured effect.−0.9pp to +1.7ppis an inconclusive test regardless of where the mean sits — Inconclusive Results - Check the lower bound against what’s worth shipping, not against zero. An interval clearing zero but sitting entirely below your break-even effect means it works and isn’t worth it — Practical vs Statistical Significance
- It depends on the prior, and at small samples it depends on it a lot. An interval reported without the prior stated is incomplete — Choosing a Prior
- It says nothing about whether the test was valid. SRM, tracking breaks and pollution all produce clean, narrow, wrong intervals — Sample Ratio Mismatch, Reading a Test Result
Why it’s the better thing to report
For communicating to people who won’t remember the definitional caveat — which is everyone, including statisticians in meetings — the credible interval is simply safer:
reported as a confidence interval → audience hears a credible interval
(and is technically wrong)
reported as a credible interval → audience hears a credible interval
(and is right)
This is a real argument, not a philosophical preference. If the sentence people will repeat is the Bayesian one either way, using the framework that licenses it removes a persistent, invisible error — Communicating Uncertainty.
Where it interacts
- Confidence Intervals — the frequentist counterpart, and the misreading this note resolves
- Prior Likelihood and Posterior — the posterior these intervals are read from
- Probability to Beat Control — the other summary of the same posterior, and a much lossier one
- Effect Size — the interval is a statement about effect size, which significance never is