Tags: statistics concept

Sampling Methods

Date: 2026-08-17


How you choose which members of a population to measure. It decides whether inference is valid at all — and the uncomfortable fact for web analytics is that you almost never choose, which makes nearly all of it a convenience sample dressed as a random one.


A sampling method is the rule for choosing which members of a population get measured; it decides whether results from the sample can be generalised to the population at all.

The methods

MethodHowBuys youCosts
Simple randomEvery member equally likelyUnbiased, simple mathsNeeds a list of the whole population
StratifiedSplit into groups, sample within eachLower variance; guarantees small groups appearNeed to know the groups in advance
ClusterSample whole groups, measure everyone in themCheap when the population is groupedHigher variance — members of a cluster are alike
SystematicEvery kth memberEasy, spreads across the listBreaks if the list has a period matching k
ConvenienceWhoever you can reachCheap, immediateBias of unknown size and direction
QuotaFill fixed counts per group, non-randomlyCheap version of stratifiedNon-random within quota — bias remains

Stratification, worked

The one method that genuinely buys precision for free, so it’s worth seeing the arithmetic.

A site with two segments behaving very differently:

segment    share    conversion rate    variance p(1−p)
mobile      70%          2.0%             0.0196
desktop     30%          5.0%             0.0475

overall conversion  =  (0.7 × 0.02) + (0.3 × 0.05)  =  0.029

Simple random sample of 10,000. The variance you face includes the between-segment differences, because the mobile/desktop mix itself varies from sample to sample.

overall variance ≈ p(1−p) = 0.029 × 0.971 = 0.02816
SE = √(0.02816 / 10,000) = 0.001678  →  0.168pp

Stratified: 7,000 mobile and 3,000 desktop, fixed by design.

SE² = Σ (share² × variance_within / n_within)

    = (0.7² × 0.0196 / 7,000) + (0.3² × 0.0475 / 3,000)
    = (0.49 × 0.0196 / 7,000)  + (0.09 × 0.0475 / 3,000)
    = 0.000001372              + 0.000001425
    = 0.000002797

SE  = √0.000002797 = 0.001672  →  0.167pp

The gain here is slight — because the segments’ variances are similar even though their rates differ. The gain from stratification grows with how different the strata are internally, and the real benefit in practice is different: it guarantees the sample contains the right proportion of each group rather than leaving it to chance.

In plain terms: stratifying stops you from accidentally drawing a sample that’s 80% mobile, which would bias the overall figure. It doesn’t usually transform your precision.

Why web analytics is a convenience sample

The honest position, and worth stating because it’s routinely glossed over:

who ends up in your data                  who is missing

visited during the period                 people who didn't visit
accepted analytics consent      ~60–80%   the rest — Consent Management
no ad blocker                   ~65–80%   blocked users
JavaScript executed                       failures, slow connections
not filtered as a bot                     misclassified real users
survived to the tracked event             those who left first

The largest single filter is usually the consent layer — Consent Management.

None of that is random selection, and the survivors differ systematically — more engaged, better connections, less privacy-conscious, more likely to convert. This is Selection Bias, and it is not fixed by having more data. A billion biased observations are exactly as biased as a thousand.

What follows practically:

  • Comparisons within the data are usually still valid, because both arms of a test are filtered the same way. This is why A/B testing survives tracking loss reasonably well — Sample Pollution
  • Absolute levels are not. “Our conversion rate is 3.2%” describes consented, unblocked, JavaScript-executing, non-bot visitors — Benchmarking
  • Segment comparisons are the dangerous ones, because filtering rates differ by segment. Safari users are filtered more aggressively than Chrome users, so a Safari-versus-Chrome conversion comparison is partly comparing filter rates — Browser Privacy Restrictions

Where you do get to choose

Sampling decisions you actually control, and where the methods above apply directly:

  • Surveys and research recruitment. Quota-sampling a panel to match your customer base is standard, and its weakness — non-random within quota — is the reason survey findings need triangulating — Surveys, Triangulation
  • Session replay review. Watching the 40 most recent replays is convenience sampling and over-weights whatever happened this morning. Stratify by segment and outcome instead — Session Replay
  • Qualitative recruitment, where the small n makes selection decisive — Sample Size in Qualitative Research
  • Vendor-side data sampling, where the tool decides for you above a threshold and you inherit its method — Data Sampling
  • Manual QA and data audits, where checking “some” orders means choosing which

The rule that survives everything

Sample size fixes noise. It never fixes bias.

biased sample, n = 1,000        estimate 4.1%, true 3.2%
biased sample, n = 1,000,000    estimate 4.1%, true 3.2%
                                ↑ the interval shrank to nothing
                                  around the wrong number

More data makes a biased estimate more confidently wrong, which is worse than being uncertain — Sampling Error, Communicating Uncertainty.

Where it interacts

  • Populations and Samples — what you’re sampling from, and why it’s usually hypothetical
  • Selection Bias — the failure mode this note exists to characterise
  • Random Variables — independence between draws is the assumption every formula rests on, and cluster sampling is the case where it fails by design
  • Randomisation Unit — the experimentation-side version of the same independence question