Tags: statistics concept
Selection Bias
Date: 2026-08-16
The sample is systematically unlike the population. It’s the error more data makes worse rather than better — and in web analytics it’s the default condition, because measurability is not randomly distributed.
What it is
Selection bias is a systematic difference between who ends up in your data and who you’re trying to describe.
SAMPLING ERROR SELECTION BIAS
random, unbiased systematic
shrinks with more data unchanged by more data
quantified by the standard quantified by nothing
error
→ a wider interval → a confidently wrong answer
That asymmetry is the whole point. A bigger biased sample is a more confidently wrong answer, and nothing in the output warns you.
Why web analytics is structurally biased
Not an implementation failure — a property of the medium. Every one of these removes a non-random slice:
| Filter | Removes |
|---|---|
| Consent denial | Privacy-conscious users, who convert differently — Consent Management |
| Ad blockers | Technical, desktop-skewed users — Ad Blockers and Tracking Loss |
| Identifier expiry | Returning users on Safari more than Chrome — Browser Privacy Restrictions |
| Delivery loss at unload | Sessions that ended abruptly — Event Batching and Delivery |
| Login requirement | Everyone who didn’t sign in — Anonymous and Identified Users |
In plain terms: the people you can measure best are not a random sample of your customers. They’re the ones whose browsers and choices happen to permit measurement, and those correlate with how they behave.
The forms it takes
- Self-selection. Anyone who chose to be in the data. Survey respondents, newsletter subscribers, logged-in users, people who left a review
- Survivorship. Only what’s still present to measure — Survivorship Bias
- Non-response. People who declined, who differ systematically from those who didn’t
- Convenience sampling. Whoever was easiest to reach. Usability testing with colleagues, or recruiting participants from your own email list
- Coverage. Your sampling frame excludes part of the population — measuring app users to conclude something about all customers
Where it produces confident errors
- “Our customers prefer X” from a survey answered by 0.4% of them, all of whom were engaged enough to respond
- Retention analysis on identified users, which excludes everyone who never returned far enough to log in — systematically overstating retention
- Conclusions from session replay, where you chose which sessions to watch — Session Replay
- Usability findings from recruited participants who are more patient and more motivated than real traffic
- Any comparison between measurable and unmeasurable populations, since only one is in the data
What doesn’t fix it
- More data. The bias is in the mechanism, not the volume
- Weighting, unless you know the direction and size of the bias — which requires knowing about the people you can’t see
- Statistical significance. A significant result from a biased sample is a precise measurement of the wrong thing
What does
Randomisation, where you control the assignment. Within an experiment, both arms are drawn from the same biased population, so the comparison is valid even though the absolute rates aren’t. This is why testing survives measurement bias that would ruin an observational analysis — Why Randomisation Works.
That’s the practical resolution worth holding: absolute numbers from web analytics are biased; differences between randomised arms are not. Which is an argument for making decisions from tests rather than from dashboards, and for treating any absolute rate as approximate.
Beyond that:
- Quantify what you’re missing. Reconcile against the order system and know the gap — Guide - Auditing a Tracking Plan
- Report the population, not just the metric. “Conversion rate (consenting users, tracked sessions)” is honest
- State the direction where you know it. Consent bias, blocking bias and survivorship all push the same way, so your measured population is more engaged than your real one