Tags: experimentation concept
Sample Pollution
Date: 2026-08-17
Users in the test who shouldn’t be, or who are counted more than once. Unlike most experimental problems it doesn’t announce itself — the test still runs, the numbers still compute, and the effect is silently diluted towards zero, which reads as “inconclusive” rather than as “broken”.
Sample pollution is any unit in a test’s data that shouldn’t be there or is counted more than once — bots, staff, users who were never exposed, users seen in both arms.
Why it biases towards null
The important property, and the reason it’s dangerous rather than merely annoying.
true effect on real users: +10% relative
test population
70,000 real users genuine +10%
30,000 bots +0% — bots don't respond to a nicer button
measured effect = (0.70 × 10%) + (0.30 × 0%) = +7%
Contamination that is unaffected by the treatment dilutes the measured effect in proportion to how much of the sample it is. 30% pollution turns a +10% effect into a +7% one — often the difference between a detected win and an inconclusive test you conclude was a failed idea.
In plain terms: pollution rarely invents a winner. It hides real ones, and it makes you believe changes don’t work — which is a much more expensive error than a false positive, because it’s invisible and it compounds across a whole programme.
The exception where it does bias away from null: if the polluting traffic is unevenly split between arms, it shifts the comparison directly. That’s a Sample Ratio Mismatch, and it’s a reason to discard the test entirely rather than to adjust for it.
The sources
| Source | What it does | Why it’s uneven |
|---|---|---|
| Bots and crawlers | Assigned, never convert | Some execute JS, some don’t — so client-side tests filter them differently by arm |
| Internal traffic | Staff testing the variant repeatedly | Staff are told to check the variant, so they’re concentrated in one arm |
| Duplicate assignment | One person counted as two | Cleared cookies, private mode, cross-device — Identity Stitching |
| Ineligible users | Assigned but couldn’t see the change | Assigned on page load, but the element is below the fold or only shows for logged-in users |
| Load testing and monitoring | Synthetic traffic hitting real endpoints | Runs on a schedule, so it clusters in time |
| Employees’ own accounts | Real purchases, unrepresentative behaviour | Staff discount patterns differ wildly |
The one that costs most: triggering on the wrong event
Not usually thought of as pollution, and usually the largest single source of it.
BAD: assign on page load
100,000 users assigned to the checkout-copy test
30,000 ever reach the step where the copy differs
70,000 could not possibly have been affected
effect diluted by 70%. a real +5% reads as +1.5%
GOOD: assign on exposure
30,000 users assigned, at the moment the differing element
is about to render
effect measured on people who could respond to it
Assign as late as possible, at the point of exposure, and no later. Later than exposure — assigning after someone interacts — is worse, because now assignment is caused by behaviour and the arms aren’t comparable at all.
The cost of late assignment is that your sample shrinks, so the test needs longer. That’s the correct trade: 30,000 exposed users beat 100,000 mostly-irrelevant ones for detecting the same effect — Experiment Assignment Tracking, Metric Selection for Tests.
Excluding, without breaking the randomisation
The rule that makes exclusions safe:
Exclude on properties that cannot have been affected by the treatment. Bot user-agent, internal IP, employee account flag, assignment timestamp — all fixed independently of which variant someone got. Excluding on those is fine and doesn’t bias anything.
Never exclude on behaviour. “Remove users who bounced immediately” or “exclude sessions under 5 seconds” filters on an outcome the treatment may have caused, which breaks the comparability randomisation bought you and can manufacture an effect from nothing — Selection Bias.
✓ exclude: user_agent matches known bot list
✓ exclude: IP in office range
✓ exclude: account flagged internal
✓ exclude: assigned before the QA sign-off timestamp
✗ exclude: session under 5 seconds ← treatment may cause short sessions
✗ exclude: users who saw an error ← treatment may cause errors
✗ exclude: "outlier" order values ← unless capped symmetrically and
declared in advance
Apply exclusions identically to both arms, and check the counts after excluding — if the exclusion removed 8% of control and 12% of variant, the exclusion itself is treatment-dependent and you’ve just created the problem you were fixing.
Catching it
- Sample Ratio Mismatch check first, always. The single most informative test of whether the sample is trustworthy, and it catches most pollution that matters
- An A test before trusting a new platform — it surfaces duplicate assignment and filtering asymmetries with no real change to confuse things
- Compare assigned counts to exposed counts. A large gap is the dilution problem above
- Watch the conversion rate of the whole test population against the site baseline. Far below it usually means bots; far above usually means you’ve excluded the wrong people
- Segment by bot-likelihood as a diagnostic, not as a filter — if excluding suspected bots changes the result materially, the result is fragile either way
Where it interacts
- Bot and Internal Traffic — the analytics-side treatment of the same population, and the two systems should agree on the definition
- Randomisation Unit — duplicate assignment is a consequence of choosing a unit you can’t identify reliably
- Experiment QA — most pollution is preventable before launch and undetectable afterwards
- Inconclusive Results — pollution is one of the standard explanations for a null, and it’s the one nobody checks