Tags: statistics concept
Sampling Error
Date: 2026-08-16
The unavoidable gap between a sample and the population it came from. It shrinks with the square root of sample size — and small samples resemble the population far less than intuition insists, which is why small segments generate so many false findings.
What it is
Sampling error is the difference between a statistic computed on a sample and the true population value, arising purely from which units happened to be sampled.
It is not a mistake. It exists in every sample, always, and no amount of care removes it — only more data reduces it.
Distinct from bias, which is a systematic difference and which more data makes worse rather than better — Selection Bias.
sampling error random, shrinks with n, quantified by the standard error
bias systematic, unaffected by n, not quantified by anything
How fast it shrinks
The square root is the constraint. Worked, on a true 3% conversion rate:
| Sample | Standard error | Rough 95% range |
|---|---|---|
| 100 | 1.71pp | 0% – 6.4% |
| 1,000 | 0.54pp | 1.9% – 4.1% |
| 10,000 | 0.17pp | 2.7% – 3.3% |
| 100,000 | 0.05pp | 2.9% – 3.1% |
At 100 visitors you cannot distinguish a 1% site from a 6% site. At 10,000 you’re within a few tenths. Every order of magnitude buys roughly a threefold improvement in precision — Standard Error.
Small samples do not resemble the population
The intuition that fails, and it fails in a specific direction: people expect small samples to be miniature versions of the population. They aren’t — they’re wildly variable.
true conversion rate 3%, samples of 100 visitors
sample A 2 conversions 2.0%
sample B 6 conversions 6.0% ← 3× sample A
sample C 1 conversion 1.0%
sample D 4 conversions 4.0%
nothing changed. Same site, same week, same 3% truth.
In plain terms: with small numbers, you’re mostly measuring which people happened to turn up. A segment that converts at twice the site average on 80 sessions is telling you nothing.
This is the mechanism behind most false findings from segmentation. Cut a population four ways and each cell has a quarter of the data and double the standard error — which is why post-hoc segments produce significance so reliably — Segmentation (test results), The Multiple Comparisons Problem.
It’s also why extremes regress: an unusually good or bad small sample got there partly by luck, and luck doesn’t repeat — Regression to the Mean.
Where it bites in practice
- Small segments in any report. Always show counts alongside rates, or a 12% conversion rate on 40 sessions reads as a finding
- Daily numbers. A single day’s conversion rate on modest traffic is mostly noise. Weekly is usually the shortest readable grain
- Low-traffic templates. A category page with 200 weekly visitors will never produce a readable rate — Percentiles in Performance on the equivalent problem in performance data
- Early test readings. The variation is widest at the start, which is exactly when people look — Peeking
- A/A tests. Two identical arms differ by sampling error alone, and 5% of the time significantly so — A-A Tests
What it isn’t
- Not a data quality problem. Clean data has it too
- Not fixed by collecting differently. Only by collecting more, or by reducing variance — Variance Reduction
- Not the same as tracking loss. Missing events are bias, not sampling error — Ad Blockers and Tracking Loss
The practical discipline
Report an interval, or report a count alongside the rate. A number with no indication of precision invites a decision the data can’t support, and the fix costs one extra column — Confidence Intervals, Communicating Uncertainty.