Tags: statistics experimentation concept

Bayesian vs Frequentist

Date: 2026-08-17


Two frames that answer different questions. The frequentist asks how surprising your data would be if nothing were happening; the Bayesian asks how likely each possible effect is, given your data. Most of the practical difference isn’t the philosophy — it’s that Bayesian output says the thing people already wrongly believe frequentist output says.


Frequentist statistics treats the true effect as fixed and asks how often the data would look like this across repeated samples; Bayesian statistics treats the effect as uncertain and updates a probability distribution over it as data arrives.

The question each answers

FREQUENTIST                          BAYESIAN

P(data this extreme | no effect)     P(effect | data observed)

"if the variant were identical       "given what I've seen, there's a
 to control, I'd see a gap this       94% probability the variant is
 big or bigger 3% of the time"        better, and the effect is most
                                      likely between +1% and +6%"

the parameter is FIXED and unknown   the parameter is a DISTRIBUTION
the DATA is random                    the data is fixed once observed

The frequentist statement is about hypothetical repetitions of the experiment. The Bayesian statement is about this experiment. That’s the entire philosophical difference, and everything below follows from it.

Why this matters practically

People misread a p-value as a Bayesian statement, universally. “p = 0.03 means there’s a 3% chance this is a fluke” is wrong, and it’s what nearly everyone hears — P-Values.

what p = 0.03 MEANS
  if there were no effect, data this extreme would occur 3% of the time

what people THINK it means
  there's a 3% chance there's no effect

these are different quantities, and they can differ enormously
— see Base Rate Fallacy

How far apart those two quantities can be is worked through in Base Rate Fallacy, and the gap is often an order of magnitude.

Bayesian output actually is the second thing. “94% probability the variant is better” means what it sounds like. This is the strongest argument for Bayesian testing in a commercial setting: not that it’s more correct, but that the number survives being repeated in a meeting by someone who didn’t run it.

Side by side

FrequentistBayesian
Outputp-value, confidence intervalPosterior distribution, credible interval, P(B>A)
Interval means”95% of intervals built this way contain the truth""95% probability the truth is in here” — Credible Intervals
Prior beliefsNot usedExplicit input — Choosing a Prior
PeekingInvalidates the result unless designed forPosterior is always valid, but see below
Decision ruleReject or fail to reject at αWhatever threshold you choose
Answers “how much”Via the intervalDirectly, as a distribution
Sample size planningWell-established, standardLess standardised
Regulatory acceptanceThe default nearly everywhereGrowing, sector-dependent

The peeking claim, corrected

The most over-sold Bayesian benefit, and it deserves precision because vendors state it loosely.

True: a posterior is a valid summary of belief given the data so far, at any moment. No correction is needed to look at it.

Also true, and usually omitted: if your rule is “stop as soon as P(B>A) > 95%”, you have a stopping rule, and stopping rules affect error rates in any framework. Repeatedly checking and stopping at the first favourable moment inflates the rate at which you ship things that don’t work — Bayesian or not.

simulate a genuinely null test (no real difference), check daily,
stop when P(B > A) > 95%

you will stop, eventually, on a large fraction of null tests
— the posterior wanders, and you're selecting on the wander

In plain terms: Bayesian methods remove the mathematical penalty for looking, not the selection problem created by choosing when to stop. What actually fixes it in both frameworks is a pre-committed stopping rule, or a design built for continuous monitoring — Peeking, Stopping Rules, Always-Valid Inference.

Where they agree

With a lot of data and a weak prior, the two converge. The posterior mean approaches the observed rate, and a 95% credible interval approaches the 95% confidence interval numerically.

conversion test, 200,000 per arm, weak prior

frequentist 95% CI on relative lift    +1.2%  to  +6.1%
Bayesian 95% credible interval         +1.2%  to  +6.1%

They’re the same numbers meaning different things. The disagreement between frameworks is loud at small samples and irrelevant at large ones — so for a well-powered test, the choice is about how you want to communicate rather than what you’ll conclude.

Choosing

Bayesian suits: communicating to non-statisticians; decisions where you want expected loss rather than a binary verdict; genuine prior information worth using; many small tests where a prior does real work; sequential decisions.

Frequentist suits: regulatory or external reporting; established sample-size planning; audiences who expect p-values; situations where you want to avoid arguing about a prior.

The honest position: for a properly powered A/B test with a pre-registered stopping rule, the framework barely matters. Effort spent choosing between them is better spent on power, metric choice and data quality, all of which dominate — Metric Selection for Tests, Statistical Power.

The dishonest use of either is the same: running until you like the number. A Bayesian tool with no stopping rule and a frequentist tool with peeking produce the same over-shipping, and neither framework protects you from it.

Where it interacts