Tags: statistics concept

Credible Intervals

Date: 2026-08-17


The interval that means what everyone wrongly thinks a confidence interval means: given the data and the prior, there’s a 95% probability the true value lies inside it. Same numbers as a confidence interval on a large sample, radically easier to state without saying something false.


A credible interval is a range that, given the data and the prior, contains the true value with a stated probability — the Bayesian counterpart of a confidence interval.

The definition people want

CREDIBLE INTERVAL (Bayesian)         CONFIDENCE INTERVAL (frequentist)

"there is a 95% probability the      "if I repeated this experiment many
 true conversion rate is between      times and built an interval each
 2.8% and 3.2%"                       time, 95% of those intervals would
                                      contain the true rate. about THIS
                                      interval, I say nothing probabilistic"

a statement about the PARAMETER      a statement about the PROCEDURE

Both sentences describe the same two numbers. The first is what people say when they report a confidence interval, and it’s wrong — the frequentist parameter is fixed, so it’s either in the interval or it isn’t, and there’s no probability about it. The Bayesian version treats the parameter as uncertain, so the probability statement is licensed — Confidence Intervals, Bayesian vs Frequentist.

Reading one off a posterior

The posterior is a distribution; the interval is the middle 95% of it.

posterior for the variant:  Beta(451, 14,551)

mean  =  451 / 15,002  =  0.03006

approximate sd of a Beta:
  √( αβ / ((α+β)²(α+β+1)) )
  = √( (451 × 14,551) / (15,002² × 15,003) )
  = √( 6,562,501 / (225,060,004 × 15,003) )
  = √( 6,562,501 / 3,376,575,240,012 )
  = √0.000001944
  = 0.001394

95% credible interval ≈ 0.03006 ± (1.96 × 0.001394)
                      =  0.02733  to  0.03279
                      →  2.73%  to  3.28%

In plain terms: given the prior and 15,000 observations, there’s a 95% probability the true conversion rate is between 2.73% and 3.28%.

Two conventions for choosing which 95%:

  • Equal-tailed — cut 2.5% off each end. What almost every tool reports, and what’s computed above
  • Highest density (HDI) — the narrowest interval containing 95%. Identical for symmetric posteriors; better for skewed ones, since an equal-tailed interval on a skewed posterior can exclude the most probable values

For conversion rates at any reasonable sample size the posterior is near-symmetric and the two coincide. For revenue metrics they can differ noticeably — Skewed and Heavy-Tailed Distributions.

The interval on the difference

What you actually want in a test isn’t each arm’s interval — it’s the interval on the difference, and this is where the Bayesian version is genuinely more convenient.

control  Beta(301, 9,701)
variant  Beta(340, 9,662)

draw 100,000 samples from each posterior, subtract pairwise,
read the percentiles of the difference

  2.5th percentile   +0.04pp
  median             +0.39pp
 97.5th percentile   +0.74pp

  →  95% credible interval on the absolute lift: +0.04pp to +0.74pp
  →  in relative terms, roughly +1.3% to +24.6%

No formula needed. With frequentist methods, the interval on a difference of proportions needs a derivation, and the interval on a ratio of proportions needs another; here you simulate and read percentiles. The same trick handles any derived quantity — expected revenue impact, difference in ratios, anything — which is the practical reason Bayesian tooling is pleasant to work with.

Reading one honestly

  • Width is the information. +0.04pp to +0.74pp is a real but imprecisely-measured effect. −0.9pp to +1.7pp is an inconclusive test regardless of where the mean sits — Inconclusive Results
  • Check the lower bound against what’s worth shipping, not against zero. An interval clearing zero but sitting entirely below your break-even effect means it works and isn’t worth it — Practical vs Statistical Significance
  • It depends on the prior, and at small samples it depends on it a lot. An interval reported without the prior stated is incomplete — Choosing a Prior
  • It says nothing about whether the test was valid. SRM, tracking breaks and pollution all produce clean, narrow, wrong intervals — Sample Ratio Mismatch, Reading a Test Result

Why it’s the better thing to report

For communicating to people who won’t remember the definitional caveat — which is everyone, including statisticians in meetings — the credible interval is simply safer:

reported as a confidence interval  →  audience hears a credible interval
                                      (and is technically wrong)

reported as a credible interval    →  audience hears a credible interval
                                      (and is right)

This is a real argument, not a philosophical preference. If the sentence people will repeat is the Bayesian one either way, using the framework that licenses it removes a persistent, invisible error — Communicating Uncertainty.

Where it interacts