Tags: statistics concept
Sampling Methods
Date: 2026-08-17
How you choose which members of a population to measure. It decides whether inference is valid at all — and the uncomfortable fact for web analytics is that you almost never choose, which makes nearly all of it a convenience sample dressed as a random one.
A sampling method is the rule for choosing which members of a population get measured; it decides whether results from the sample can be generalised to the population at all.
The methods
| Method | How | Buys you | Costs |
|---|---|---|---|
| Simple random | Every member equally likely | Unbiased, simple maths | Needs a list of the whole population |
| Stratified | Split into groups, sample within each | Lower variance; guarantees small groups appear | Need to know the groups in advance |
| Cluster | Sample whole groups, measure everyone in them | Cheap when the population is grouped | Higher variance — members of a cluster are alike |
| Systematic | Every kth member | Easy, spreads across the list | Breaks if the list has a period matching k |
| Convenience | Whoever you can reach | Cheap, immediate | Bias of unknown size and direction |
| Quota | Fill fixed counts per group, non-randomly | Cheap version of stratified | Non-random within quota — bias remains |
Stratification, worked
The one method that genuinely buys precision for free, so it’s worth seeing the arithmetic.
A site with two segments behaving very differently:
segment share conversion rate variance p(1−p)
mobile 70% 2.0% 0.0196
desktop 30% 5.0% 0.0475
overall conversion = (0.7 × 0.02) + (0.3 × 0.05) = 0.029
Simple random sample of 10,000. The variance you face includes the between-segment differences, because the mobile/desktop mix itself varies from sample to sample.
overall variance ≈ p(1−p) = 0.029 × 0.971 = 0.02816
SE = √(0.02816 / 10,000) = 0.001678 → 0.168pp
Stratified: 7,000 mobile and 3,000 desktop, fixed by design.
SE² = Σ (share² × variance_within / n_within)
= (0.7² × 0.0196 / 7,000) + (0.3² × 0.0475 / 3,000)
= (0.49 × 0.0196 / 7,000) + (0.09 × 0.0475 / 3,000)
= 0.000001372 + 0.000001425
= 0.000002797
SE = √0.000002797 = 0.001672 → 0.167pp
The gain here is slight — because the segments’ variances are similar even though their rates differ. The gain from stratification grows with how different the strata are internally, and the real benefit in practice is different: it guarantees the sample contains the right proportion of each group rather than leaving it to chance.
In plain terms: stratifying stops you from accidentally drawing a sample that’s 80% mobile, which would bias the overall figure. It doesn’t usually transform your precision.
Why web analytics is a convenience sample
The honest position, and worth stating because it’s routinely glossed over:
who ends up in your data who is missing
visited during the period people who didn't visit
accepted analytics consent ~60–80% the rest — Consent Management
no ad blocker ~65–80% blocked users
JavaScript executed failures, slow connections
not filtered as a bot misclassified real users
survived to the tracked event those who left first
The largest single filter is usually the consent layer — Consent Management.
None of that is random selection, and the survivors differ systematically — more engaged, better connections, less privacy-conscious, more likely to convert. This is Selection Bias, and it is not fixed by having more data. A billion biased observations are exactly as biased as a thousand.
What follows practically:
- Comparisons within the data are usually still valid, because both arms of a test are filtered the same way. This is why A/B testing survives tracking loss reasonably well — Sample Pollution
- Absolute levels are not. “Our conversion rate is 3.2%” describes consented, unblocked, JavaScript-executing, non-bot visitors — Benchmarking
- Segment comparisons are the dangerous ones, because filtering rates differ by segment. Safari users are filtered more aggressively than Chrome users, so a Safari-versus-Chrome conversion comparison is partly comparing filter rates — Browser Privacy Restrictions
Where you do get to choose
Sampling decisions you actually control, and where the methods above apply directly:
- Surveys and research recruitment. Quota-sampling a panel to match your customer base is standard, and its weakness — non-random within quota — is the reason survey findings need triangulating — Surveys, Triangulation
- Session replay review. Watching the 40 most recent replays is convenience sampling and over-weights whatever happened this morning. Stratify by segment and outcome instead — Session Replay
- Qualitative recruitment, where the small n makes selection decisive — Sample Size in Qualitative Research
- Vendor-side data sampling, where the tool decides for you above a threshold and you inherit its method — Data Sampling
- Manual QA and data audits, where checking “some” orders means choosing which
The rule that survives everything
Sample size fixes noise. It never fixes bias.
biased sample, n = 1,000 estimate 4.1%, true 3.2%
biased sample, n = 1,000,000 estimate 4.1%, true 3.2%
↑ the interval shrank to nothing
around the wrong number
More data makes a biased estimate more confidently wrong, which is worse than being uncertain — Sampling Error, Communicating Uncertainty.
Where it interacts
- Populations and Samples — what you’re sampling from, and why it’s usually hypothetical
- Selection Bias — the failure mode this note exists to characterise
- Random Variables — independence between draws is the assumption every formula rests on, and cluster sampling is the case where it fails by design
- Randomisation Unit — the experimentation-side version of the same independence question