Tags: statistics concept

Populations and Samples

Date: 2026-08-16


You never measure the population. Every number you have is a sample statistic standing in for something you can’t see — which is why uncertainty is unavoidable rather than a sign you collected badly.


The two things

  • Population — every unit you want to make a claim about. All visitors to the site, including next month’s
  • Sample — the subset you actually measured

A number computed on the population is a parameter. The same number computed on a sample is a statistic. Conventionally: population mean μ, sample mean x̄; population standard deviation σ, sample standard deviation s.

POPULATION  (unobservable)          SAMPLE  (what you have)

all visitors, past and future       last 30 days, 51,428 visitors
true conversion rate  μ = ?         observed rate  x̄ = 3.02%
true spread           σ = ?         observed spread s = £13.44

The statistic is an estimate of the parameter. That gap is what Standard Error, Confidence Intervals and every hypothesis test exist to quantify.

”We have all the data” is the field’s biggest error

The most common objection in analytics: we don’t need statistics, we have every visitor, this is the population not a sample.

It isn’t, and the reason matters.

  • The population includes the future. You aren’t asking “did these 51,428 people convert at 3.02%” — you know that, exactly, and it’s not a decision-relevant question. You’re asking “will visitors convert better if we ship this”, and next month’s visitors are in the population but not in your data
  • The population is the process, not the records. Your site converts at some underlying rate; last month’s traffic is one realisation of it. Run the same month again and you’d get a different number without anything changing
  • Which means every observed difference contains noise, and telling noise from signal is what the rest of this domain is for

In plain terms: having every row in the database tells you what happened. It doesn’t tell you what will happen next time, and every decision is about next time.

The narrow exception: a purely historical, backward-looking question — “how much revenue did we take in July” — genuinely is a population question, and needs no inference. The moment you compare, forecast or decide, you’re back to sampling.

What the sample has to be

A sample supports claims about a population only if it’s representative — every member of the population had a known, non-zero chance of being in it.

Ways this fails silently, all common:

  • Only measurable users are measured. Consent denial and ad blockers remove a non-random slice — see Ad Blockers and Tracking Loss
  • The sample is a time window that isn’t typical. A fortnight containing payday, a bank holiday or a promotion
  • The population changed underneath you. A campaign shifted the traffic mix, so last month’s sample describes a different population than this month’s

None of these are fixed by collecting more. A bigger biased sample is a more confidently wrong answer — see Selection Bias.

Why this note comes first

Everything downstream is an answer to “how far might the statistic be from the parameter”:

Get the population wrong — measure UK visitors and claim something about all visitors — and none of the machinery helps. It quantifies sampling error, not the wrong question.