Tags: statistics concept

Choosing a Prior

Date: 2026-08-17


Where the subjectivity in Bayesian analysis lives. “Uninformative” priors are not neutral — a uniform prior on a conversion rate asserts that 50% is as plausible as 3%, which is a strong and false belief — and the honest move is to state a weak, sensible prior rather than to pretend you have none.


A prior is the probability distribution over the true value that a Bayesian analysis starts from, before seeing the data; choosing a prior is deciding what that starting belief is and how strongly it’s held.

Why “uninformative” isn’t

Beta(1, 1) — the uniform prior, often labelled "uninformative"

says: P(true rate is between 0% and 10%)   =  10%
      P(true rate is between 40% and 50%)  =  10%
      P(true rate is above 50%)            =  50%

for an ecommerce conversion rate, that last line is
a claim that a coin flip is the most likely outcome.
that isn't neutrality. it's a very strong wrong belief.

Every prior is a claim. Refusing to choose one just means accepting whatever the tool’s default asserts — and the default is usually uniform, which is the least plausible option available for a rate you already know is small.

The two things a prior encodes

Separate them, because they’re chosen differently:

LOCATION    where you think the value is        →  the mean, α/(α+β)
STRENGTH    how much evidence that belief       →  α + β, in
            is worth                                imaginary observations

Strength is the one that matters and the one people neglect. Get the location roughly right and the strength deliberately low, and the prior does its job — mild scepticism about extreme results — without overriding your data.

belief: "our conversion rate is around 3%"

Beta(3, 97)         mean 3%,  strength 100      weak. sensible default
Beta(30, 970)       mean 3%,  strength 1,000    moderate
Beta(300, 9,700)    mean 3%,  strength 10,000   strong. only if you
                                                 genuinely have a year
                                                 of stable history

Rule of thumb: the prior’s strength should be well below your expected sample size, so the data wins. If you’re collecting 200,000 observations, a strength of 100–1,000 expresses a belief without dominating anything — Prior Likelihood and Posterior.

The prior that matters most: on the effect, not the rate

This is the part usually skipped, and it’s where a prior earns its keep.

Priors on each arm’s rate barely matter at scale. The consequential prior is on the difference — and here you have real information, because you know what test results look like at your organisation.

what your archive says about effect sizes

  ~70% of tests: effect within ±1% relative
  ~25%:          ±1% to ±5%
   ~5%:          beyond ±5%
  almost none:   beyond ±20%

a prior centred on ZERO with most mass within ±5% is a
faithful description of reality — Experiment Archive

That distribution is sitting in your own records, unused — Experiment Archive.

A sceptical prior on the effect is the single most effective correction for Winner’s Curse. It shrinks extreme observed results towards zero, and extreme observed results are precisely the ones most inflated by selection:

observed relative lift  +40%,  from a small test
weak prior              posterior ≈ +36%       ← barely shrunk
sceptical prior         posterior ≈ +8%        ← shrunk hard, and much
                                                 closer to what re-tests
                                                 of such results deliver

In plain terms: if you know from experience that +40% effects essentially never happen, a method that reports +40% because one small test said so is ignoring what you know. The sceptical prior is that knowledge, written down.

How much it actually moves the answer

The reassuring arithmetic, and the reason not to over-worry:

data: 300 conversions in 10,000       observed 3.000%

prior                posterior mean      shift from observed
Beta(1, 1)              3.009%             +0.009pp
Beta(3, 97)             3.000%              0.000pp
Beta(30, 970)           3.000%              0.000pp
Beta(300, 9,700)        3.000%              0.000pp

  → at 10,000 observations, with the prior centred correctly,
    the choice is irrelevant

now the same priors with only 200 observations, 12 conversions (6%)

Beta(1, 1)              6.44%
Beta(3, 97)             5.00%
Beta(30, 970)           3.59%
Beta(300, 9,700)        3.06%

  → at 200 observations the prior decides the answer

The prior matters exactly when your data is thin — which is exactly when you should be sceptical anyway. That’s the mechanism working, not a flaw.

Practical rules

  • Never use a uniform prior on a rate you know is small. Beta(3, 97) costs nothing and is honest
  • Centre the effect prior on zero. Most changes do nothing; a prior saying otherwise is optimism encoded as maths
  • Set strength from your archive, not from feel. The distribution of your own past effect sizes is the best-justified prior available and almost nobody uses it
  • Never choose a prior after seeing the data. It’s the Bayesian form of P-Hacking, and it’s undetectable in the write-up
  • State the prior when reporting. A posterior without its prior is uninterpretable, and most tool interfaces hide it
  • Do a sensitivity check. Re-run with a weak, a moderate and a sceptical prior. If the decision flips, say so — the data isn’t strong enough to decide on its own, which is itself the finding
  • Never let a prior encode what you want. “We’re confident this will work, so a favourable prior” is assuming the conclusion

The critique, stated fairly

The frequentist objection is that the prior lets the analyst influence the result, and that’s true. Two honest responses:

It’s visible. A prior is one stated assumption that can be challenged and varied. Frequentist analysis carries equivalent subjectivity — metric choice, stopping rules, exclusions, model form, which comparisons get reported — spread across a dozen decisions that never appear in the write-up. Concentrating it in one auditable place is arguably an improvement.

It converges. With adequate data and a sensibly-centred prior, the choice stops mattering, as the table above shows. The disagreement is loud precisely where the evidence is weak, which is where disagreement belongs.

Where it interacts