Tags: statistics concept

Prior Likelihood and Posterior

Date: 2026-08-17


The three parts of a Bayesian update: what you believed, what the data says, and what you believe afterwards. For conversion rates the arithmetic is unusually friendly — the update reduces to adding your conversions and non-conversions to two numbers.


The prior is the belief about a value before the data; the likelihood is how probable the observed data is under each possible value; the posterior is the updated belief, prior and likelihood combined.

The rule

posterior  ∝  likelihood  ×  prior

P(effect | data)  ∝  P(data | effect)  ×  P(effect)
  • Prior — what you believed about the conversion rate before this test
  • Likelihood — how probable the observed data is, for each possible rate
  • Posterior — the updated belief, which becomes the prior for next time

The ∝ (“proportional to”) hides a normalising constant that makes it sum to 1. It rarely needs computing, because the conjugate case below gives the answer directly.

The conjugate shortcut

For yes/no outcomes there’s a distribution — the Beta distribution, written Beta(α, β) — with a convenient property: if your prior is Beta and your data is conversions out of trials, the posterior is also Beta, and you get it by addition.

prior       Beta(α, β)
data        s conversions out of n trials
posterior   Beta(α + s,  β + n − s)

Read Beta(α, β) as “α prior conversions and β prior non-conversions”. That’s the whole intuition — a prior is just imaginary data you’re bringing along.

mean of Beta(α, β)  =  α / (α + β)

Worked, with a weak prior

PRIOR      Beta(1, 1)          uniform — every rate 0% to 100% equally likely
           mean = 1/2 = 0.500  ← already an absurd belief for a
                                  conversion rate. see Choosing a Prior

DATA       300 conversions in 10,000 sessions

POSTERIOR  Beta(1 + 300,  1 + 10,000 − 300)
        =  Beta(301, 9,701)

           mean  =  301 / (301 + 9,701)
                 =  301 / 10,002
                 =  0.03009   →  3.009%

The observed rate was 300/10,000 = 3.000%, and the posterior mean is 3.009%. The prior pulled it up by nine thousandths of a percentage point — negligible, because one imaginary conversion against 10,000 real ones is nothing.

In plain terms: with enough data, the prior stops mattering. That’s the reassuring property, and it’s why arguments about priors matter far less on high-traffic tests than they sound like they should.

Worked, with a strong prior and little data

Now the case where it does matter.

PRIOR      Beta(300, 9,700)     "we've seen ~3% across a year of data"
           mean = 300/10,000 = 0.0300

DATA       12 conversions in 200 sessions     observed rate = 6.0%

POSTERIOR  Beta(300 + 12,  9,700 + 200 − 12)
        =  Beta(312, 9,888)

           mean  =  312 / 10,200  =  0.03059  →  3.06%

The data said 6%. The posterior says 3.06%. The prior — carrying the weight of 10,000 prior observations — almost entirely overwhelms 200 new ones, and it should: 12 conversions from 200 sessions is completely consistent with a true rate of 3%.

P(observing ≥12 conversions in 200 | true rate 3%)
expected = 6, sd = √(200 × 0.03 × 0.97) = 2.41
12 is about 2.5 sd above expectation — unusual, not extraordinary

This is the mechanism doing exactly what it should: refusing to be impressed by a small sample. A frequentist test on the same data would report a large observed lift with a very wide interval; the Bayesian version folds the prior in and reports a modest posterior. Same scepticism, expressed differently.

The weight of the prior, made explicit

α + β is the prior’s strength, measured in imaginary observations.

prior            α + β      equivalent to        pull on 1,000 new observations

Beta(1, 1)          2       2 observations       essentially none
Beta(3, 97)       100       100 observations     slight
Beta(30, 970)   1,000       1,000 observations   equal weight to your data
Beta(300, 9700) 10,000      10,000 observations  dominates 1,000 new ones

Choose α + β to state how much evidence your prior belief is worth. If you’d want a year of history to be outweighed by two weeks of test data, the prior’s strength should be smaller than the test’s sample. This is a far more useful way to set a prior than arguing about distributions — Choosing a Prior.

Sequential updating

Posteriors chain, and the order doesn’t matter:

Beta(1,1)  →  +300/10,000  →  Beta(301, 9701)
           →  +150/5,000   →  Beta(451, 14,551)

same as updating once with 450/15,000:
Beta(1 + 450, 1 + 15,000 − 450) = Beta(451, 14,551)   ✓

This is what makes Bayesian methods feel natural for continuous monitoring — the posterior after each batch is a complete, valid summary. It does not make an “stop when it looks good” rule safe, which is a separate issue — Bayesian vs Frequentist, Stopping Rules.

What it doesn’t do

  • It doesn’t remove subjectivity — it relocates it. The prior is a choice, made by a person, that changes the answer. Frequentist analysis has the same subjectivity distributed across metric choice, stopping rules and model form; Bayesian methods concentrate it in one visible place, which is arguably better and definitely more arguable
  • The conjugate shortcut is specific to yes/no data. Revenue per visitor has no such tidy form and needs simulation — which is what tools do underneath
  • Garbage data still gives a confident posterior. A broken test with SRM produces a beautifully tight posterior around the wrong number

Where it interacts