Tags: statistics concept
Prior Likelihood and Posterior
Date: 2026-08-17
The three parts of a Bayesian update: what you believed, what the data says, and what you believe afterwards. For conversion rates the arithmetic is unusually friendly — the update reduces to adding your conversions and non-conversions to two numbers.
The prior is the belief about a value before the data; the likelihood is how probable the observed data is under each possible value; the posterior is the updated belief, prior and likelihood combined.
The rule
posterior ∝ likelihood × prior
P(effect | data) ∝ P(data | effect) × P(effect)
- Prior — what you believed about the conversion rate before this test
- Likelihood — how probable the observed data is, for each possible rate
- Posterior — the updated belief, which becomes the prior for next time
The ∝ (“proportional to”) hides a normalising constant that makes it sum to 1. It rarely needs computing, because the conjugate case below gives the answer directly.
The conjugate shortcut
For yes/no outcomes there’s a distribution — the Beta distribution, written Beta(α, β) — with a convenient property: if your prior is Beta and your data is conversions out of trials, the posterior is also Beta, and you get it by addition.
prior Beta(α, β)
data s conversions out of n trials
posterior Beta(α + s, β + n − s)
Read Beta(α, β) as “α prior conversions and β prior non-conversions”. That’s the whole intuition — a prior is just imaginary data you’re bringing along.
mean of Beta(α, β) = α / (α + β)
Worked, with a weak prior
PRIOR Beta(1, 1) uniform — every rate 0% to 100% equally likely
mean = 1/2 = 0.500 ← already an absurd belief for a
conversion rate. see Choosing a Prior
DATA 300 conversions in 10,000 sessions
POSTERIOR Beta(1 + 300, 1 + 10,000 − 300)
= Beta(301, 9,701)
mean = 301 / (301 + 9,701)
= 301 / 10,002
= 0.03009 → 3.009%
The observed rate was 300/10,000 = 3.000%, and the posterior mean is 3.009%. The prior pulled it up by nine thousandths of a percentage point — negligible, because one imaginary conversion against 10,000 real ones is nothing.
In plain terms: with enough data, the prior stops mattering. That’s the reassuring property, and it’s why arguments about priors matter far less on high-traffic tests than they sound like they should.
Worked, with a strong prior and little data
Now the case where it does matter.
PRIOR Beta(300, 9,700) "we've seen ~3% across a year of data"
mean = 300/10,000 = 0.0300
DATA 12 conversions in 200 sessions observed rate = 6.0%
POSTERIOR Beta(300 + 12, 9,700 + 200 − 12)
= Beta(312, 9,888)
mean = 312 / 10,200 = 0.03059 → 3.06%
The data said 6%. The posterior says 3.06%. The prior — carrying the weight of 10,000 prior observations — almost entirely overwhelms 200 new ones, and it should: 12 conversions from 200 sessions is completely consistent with a true rate of 3%.
P(observing ≥12 conversions in 200 | true rate 3%)
expected = 6, sd = √(200 × 0.03 × 0.97) = 2.41
12 is about 2.5 sd above expectation — unusual, not extraordinary
This is the mechanism doing exactly what it should: refusing to be impressed by a small sample. A frequentist test on the same data would report a large observed lift with a very wide interval; the Bayesian version folds the prior in and reports a modest posterior. Same scepticism, expressed differently.
The weight of the prior, made explicit
α + β is the prior’s strength, measured in imaginary observations.
prior α + β equivalent to pull on 1,000 new observations
Beta(1, 1) 2 2 observations essentially none
Beta(3, 97) 100 100 observations slight
Beta(30, 970) 1,000 1,000 observations equal weight to your data
Beta(300, 9700) 10,000 10,000 observations dominates 1,000 new ones
Choose α + β to state how much evidence your prior belief is worth. If you’d want a year of history to be outweighed by two weeks of test data, the prior’s strength should be smaller than the test’s sample. This is a far more useful way to set a prior than arguing about distributions — Choosing a Prior.
Sequential updating
Posteriors chain, and the order doesn’t matter:
Beta(1,1) → +300/10,000 → Beta(301, 9701)
→ +150/5,000 → Beta(451, 14,551)
same as updating once with 450/15,000:
Beta(1 + 450, 1 + 15,000 − 450) = Beta(451, 14,551) ✓
This is what makes Bayesian methods feel natural for continuous monitoring — the posterior after each batch is a complete, valid summary. It does not make an “stop when it looks good” rule safe, which is a separate issue — Bayesian vs Frequentist, Stopping Rules.
What it doesn’t do
- It doesn’t remove subjectivity — it relocates it. The prior is a choice, made by a person, that changes the answer. Frequentist analysis has the same subjectivity distributed across metric choice, stopping rules and model form; Bayesian methods concentrate it in one visible place, which is arguably better and definitely more arguable
- The conjugate shortcut is specific to yes/no data. Revenue per visitor has no such tidy form and needs simulation — which is what tools do underneath
- Garbage data still gives a confident posterior. A broken test with SRM produces a beautifully tight posterior around the wrong number
Where it interacts
- Credible Intervals — reading a range off the posterior
- Probability to Beat Control — comparing two posteriors, which is what a Bayesian A/B tool reports
- Base Rate Fallacy — the same update, applied to diagnosis rather than to rates, and the clearest demonstration of why priors can’t be ignored
- Binomial and Bernoulli Distributions — the likelihood half of the update