Tags: statistics experimentation concept
Bayesian vs Frequentist
Date: 2026-08-17
Two frames that answer different questions. The frequentist asks how surprising your data would be if nothing were happening; the Bayesian asks how likely each possible effect is, given your data. Most of the practical difference isn’t the philosophy — it’s that Bayesian output says the thing people already wrongly believe frequentist output says.
Frequentist statistics treats the true effect as fixed and asks how often the data would look like this across repeated samples; Bayesian statistics treats the effect as uncertain and updates a probability distribution over it as data arrives.
The question each answers
FREQUENTIST BAYESIAN
P(data this extreme | no effect) P(effect | data observed)
"if the variant were identical "given what I've seen, there's a
to control, I'd see a gap this 94% probability the variant is
big or bigger 3% of the time" better, and the effect is most
likely between +1% and +6%"
the parameter is FIXED and unknown the parameter is a DISTRIBUTION
the DATA is random the data is fixed once observed
The frequentist statement is about hypothetical repetitions of the experiment. The Bayesian statement is about this experiment. That’s the entire philosophical difference, and everything below follows from it.
Why this matters practically
People misread a p-value as a Bayesian statement, universally. “p = 0.03 means there’s a 3% chance this is a fluke” is wrong, and it’s what nearly everyone hears — P-Values.
what p = 0.03 MEANS
if there were no effect, data this extreme would occur 3% of the time
what people THINK it means
there's a 3% chance there's no effect
these are different quantities, and they can differ enormously
— see Base Rate Fallacy
How far apart those two quantities can be is worked through in Base Rate Fallacy, and the gap is often an order of magnitude.
Bayesian output actually is the second thing. “94% probability the variant is better” means what it sounds like. This is the strongest argument for Bayesian testing in a commercial setting: not that it’s more correct, but that the number survives being repeated in a meeting by someone who didn’t run it.
Side by side
| Frequentist | Bayesian | |
|---|---|---|
| Output | p-value, confidence interval | Posterior distribution, credible interval, P(B>A) |
| Interval means | ”95% of intervals built this way contain the truth" | "95% probability the truth is in here” — Credible Intervals |
| Prior beliefs | Not used | Explicit input — Choosing a Prior |
| Peeking | Invalidates the result unless designed for | Posterior is always valid, but see below |
| Decision rule | Reject or fail to reject at α | Whatever threshold you choose |
| Answers “how much” | Via the interval | Directly, as a distribution |
| Sample size planning | Well-established, standard | Less standardised |
| Regulatory acceptance | The default nearly everywhere | Growing, sector-dependent |
The peeking claim, corrected
The most over-sold Bayesian benefit, and it deserves precision because vendors state it loosely.
True: a posterior is a valid summary of belief given the data so far, at any moment. No correction is needed to look at it.
Also true, and usually omitted: if your rule is “stop as soon as P(B>A) > 95%”, you have a stopping rule, and stopping rules affect error rates in any framework. Repeatedly checking and stopping at the first favourable moment inflates the rate at which you ship things that don’t work — Bayesian or not.
simulate a genuinely null test (no real difference), check daily,
stop when P(B > A) > 95%
you will stop, eventually, on a large fraction of null tests
— the posterior wanders, and you're selecting on the wander
In plain terms: Bayesian methods remove the mathematical penalty for looking, not the selection problem created by choosing when to stop. What actually fixes it in both frameworks is a pre-committed stopping rule, or a design built for continuous monitoring — Peeking, Stopping Rules, Always-Valid Inference.
Where they agree
With a lot of data and a weak prior, the two converge. The posterior mean approaches the observed rate, and a 95% credible interval approaches the 95% confidence interval numerically.
conversion test, 200,000 per arm, weak prior
frequentist 95% CI on relative lift +1.2% to +6.1%
Bayesian 95% credible interval +1.2% to +6.1%
They’re the same numbers meaning different things. The disagreement between frameworks is loud at small samples and irrelevant at large ones — so for a well-powered test, the choice is about how you want to communicate rather than what you’ll conclude.
Choosing
Bayesian suits: communicating to non-statisticians; decisions where you want expected loss rather than a binary verdict; genuine prior information worth using; many small tests where a prior does real work; sequential decisions.
Frequentist suits: regulatory or external reporting; established sample-size planning; audiences who expect p-values; situations where you want to avoid arguing about a prior.
The honest position: for a properly powered A/B test with a pre-registered stopping rule, the framework barely matters. Effort spent choosing between them is better spent on power, metric choice and data quality, all of which dominate — Metric Selection for Tests, Statistical Power.
The dishonest use of either is the same: running until you like the number. A Bayesian tool with no stopping rule and a frequentist tool with peeking produce the same over-shipping, and neither framework protects you from it.
Where it interacts
- Prior Likelihood and Posterior — the mechanism, in one worked update
- Probability to Beat Control — the output most Bayesian test tools report, and its specific trap
- P-Values — the frequentist output and the four things it isn’t
- Confidence Intervals — the frequentist interval, and why its definition is so much harder to state than a credible interval’s