Tags: statistics concept
Binomial and Bernoulli Distributions
Date: 2026-08-16
The maths of yes/no outcomes, which is every conversion rate you’ll ever test. Its variance formula is why low-converting sites are so expensive to test on, and that single fact explains more about testing budgets than anything else.
What they are
A Bernoulli trial is one yes/no event with probability p of success. One visitor: converted or didn’t.
A binomial distribution is the count of successes across n independent Bernoulli trials. Out of 53,000 visitors, how many converted.
Bernoulli one visitor → 0 or 1
Binomial 53,000 visitors → a count, distributed around 1,590
The variance formula
The whole practical payoff, and it’s one line:
Worked, per visitor:
| Conversion rate | p(1−p) | Standard deviation |
|---|---|---|
| 1% | 0.0099 | 0.0995 |
| 3% | 0.0291 | 0.1706 |
| 10% | 0.0900 | 0.3000 |
| 50% | 0.2500 | 0.5000 |
Variance is highest at 50% and falls towards either extreme. That’s intuitive — a coin flip is maximally uncertain, a 1-in-100 event is nearly always “no”.
Why low-converting sites are expensive to test
The counterintuitive part, and it matters commercially.
Absolute variance is lower at 1% than at 10%. But you test relative lifts, and the effect you’re detecting shrinks faster than the variance does.
detect a 10% relative lift, 80% power, 95% significance
10% baseline → 11% δ = 0.010 n ≈ 15,000 per arm
3% baseline → 3.3% δ = 0.003 n ≈ 53,000
1% baseline → 1.1% δ = 0.001 n ≈ 163,000
A 1% site needs roughly three times the traffic of a 3% site to answer the same question. Nothing about the site is worse — the arithmetic is just harsher at low base rates.
In plain terms: the rarer the thing you’re measuring, the more people you need to see enough of it. Low-converting businesses aren’t just harder commercially, they’re harder to learn about.
This is also why upper-funnel metrics are cheap to test on: add-to-cart at 12% needs a fraction of the traffic that purchase at 3% does — Metric Sensitivity.
Where it feeds
- Sample Size Calculation — the
2p̄(1−p̄)term in the formula is this variance, doubled for two arms - Proportion Tests — the standard error of a proportion is √(p(1−p)/n), straight from here
- Minimum Detectable Effect — the square relationship between effect and sample size, combined with this, is why halving your MDE quadruples the traffic
The approximation
At scale, a binomial distribution is closely approximated by a normal one, which is what makes standard tests applicable. The condition is roughly 10 expected successes and 10 expected failures per arm.
At a 0.5% conversion rate that’s 2,000 users per arm before the approximation is even reasonable, and far more before the test is powered. For genuinely rare events — complaints, chargebacks, returns above a threshold — the approximation fails and an exact method or Bootstrapping is the honest route.
The independence assumption
Both distributions assume trials are independent. Every conversion test relies on this, and it’s the assumption most easily broken:
- The same user counted twice — sessions as the unit rather than users. Their two sessions aren’t independent, which understates variance and overstates significance — Randomisation Unit, Sessionisation
- Shared influence — users on the same account, or affected by the same stock-out
- Time correlation — a promotion affecting everyone in one period
Violating independence makes results look more significant than they are, which is the dangerous direction and produces no warning.