Data Sampling
Date: 2026-08-16
Your tool stopped counting everything and estimated the rest. It’s usually fine for a headline number and quietly destroys small segments — and most tools indicate it somewhere you’re not looking.
What it is
Sampling in analytics is a tool computing a result from a subset of events and scaling up, rather than processing all of them. It’s a cost control: full computation over billions of rows is expensive, and an estimate is nearly as good for most questions.
Two kinds, and they behave differently:
- Query-time sampling — the tool processes a fraction when you ask a complex question. Varies by query, and by date range
- Collection-time sampling — only a fraction of events are ever recorded. Permanent, and no amount of querying recovers the rest
Why it hurts segments specifically
The headline number survives sampling well. Anything narrow does not.
1,000,000 sessions, 10% sample = 100,000 counted
site conversion rate 3.0% ← fine, ±0.1pp
mobile, paid social 3,000 sampled sessions ← still readable
mobile, paid social,
Scotland, size guide 12 sampled sessions
→ scaled to 120
→ one more or fewer in the sample
moves the reported figure by 10
In plain terms: sampling multiplies whatever it found. Where it found almost nothing, it multiplies almost nothing, and the answer swings wildly on individual events.
The cruel part is that the deeper you segment — which is where the insight is — the worse it gets, and the number still renders confidently with no warning attached.
Where it appears
- Analytics tools at volume, typically on unsampled thresholds tied to a paid tier [CHECK: current thresholds and which report types sample, per tool — these change and are tier-dependent]
- Session replay — almost always sampled, and often not randomly, which biases what you watch — Session Replay
- Real user monitoring, deliberately, to control volume — Real User Monitoring
- Your own collection, if someone set a sample rate to control cost
Detecting it
- Look for the indicator. Most tools show a sampling notice, usually small, at the top of a report. Check it before quoting any number
- Re-run the same query twice. Different answers means query-time sampling
- Narrow the date range. If a figure changes disproportionately, you crossed a sampling threshold
- Check absolute counts, not just rates. Suspiciously round numbers — exactly 4,200 — suggest scaling
Living with it
- Shorten the date range to get under the threshold, then combine periods yourself
- Simplify the query. Fewer dimensions and no custom segments often avoids triggering it
- Use the raw export. A warehouse export is unsampled, which is one of the strongest practical arguments for Warehouse-First Analytics — and it’s the only fix that works for the deep segments where sampling does most damage
- Never compare a sampled figure to an unsampled one. They’re different measurements, and this is a common source of apparent Tool Discrepancies
- Never report a sampled segment count as though it were a count. Report it as an estimate, or don’t report it
Sampling isn’t the enemy
Worth being clear: a random sample is a legitimate statistical instrument, and a 10% sample of a million sessions is more than enough to estimate a site-wide rate precisely — see Sampling Error.
The problems are specific: it isn’t disclosed prominently, it interacts badly with fine segmentation, and it’s applied by the tool rather than chosen by you. A sample you designed with a known rate is fine. A sample the tool applied silently, at a rate that varies by query, is a number you can’t reason about.