Tags: experimentation concept

Sample Pollution

Date: 2026-08-17


Users in the test who shouldn’t be, or who are counted more than once. Unlike most experimental problems it doesn’t announce itself — the test still runs, the numbers still compute, and the effect is silently diluted towards zero, which reads as “inconclusive” rather than as “broken”.


Sample pollution is any unit in a test’s data that shouldn’t be there or is counted more than once — bots, staff, users who were never exposed, users seen in both arms.

Why it biases towards null

The important property, and the reason it’s dangerous rather than merely annoying.

true effect on real users:  +10% relative

test population
  70,000 real users        genuine +10%
  30,000 bots              +0% — bots don't respond to a nicer button

measured effect  =  (0.70 × 10%) + (0.30 × 0%)  =  +7%

Contamination that is unaffected by the treatment dilutes the measured effect in proportion to how much of the sample it is. 30% pollution turns a +10% effect into a +7% one — often the difference between a detected win and an inconclusive test you conclude was a failed idea.

In plain terms: pollution rarely invents a winner. It hides real ones, and it makes you believe changes don’t work — which is a much more expensive error than a false positive, because it’s invisible and it compounds across a whole programme.

The exception where it does bias away from null: if the polluting traffic is unevenly split between arms, it shifts the comparison directly. That’s a Sample Ratio Mismatch, and it’s a reason to discard the test entirely rather than to adjust for it.

The sources

SourceWhat it doesWhy it’s uneven
Bots and crawlersAssigned, never convertSome execute JS, some don’t — so client-side tests filter them differently by arm
Internal trafficStaff testing the variant repeatedlyStaff are told to check the variant, so they’re concentrated in one arm
Duplicate assignmentOne person counted as twoCleared cookies, private mode, cross-device — Identity Stitching
Ineligible usersAssigned but couldn’t see the changeAssigned on page load, but the element is below the fold or only shows for logged-in users
Load testing and monitoringSynthetic traffic hitting real endpointsRuns on a schedule, so it clusters in time
Employees’ own accountsReal purchases, unrepresentative behaviourStaff discount patterns differ wildly

The one that costs most: triggering on the wrong event

Not usually thought of as pollution, and usually the largest single source of it.

BAD: assign on page load

  100,000 users assigned to the checkout-copy test
   30,000 ever reach the step where the copy differs
   70,000 could not possibly have been affected

  effect diluted by 70%. a real +5% reads as +1.5%

GOOD: assign on exposure

  30,000 users assigned, at the moment the differing element
  is about to render

  effect measured on people who could respond to it

Assign as late as possible, at the point of exposure, and no later. Later than exposure — assigning after someone interacts — is worse, because now assignment is caused by behaviour and the arms aren’t comparable at all.

The cost of late assignment is that your sample shrinks, so the test needs longer. That’s the correct trade: 30,000 exposed users beat 100,000 mostly-irrelevant ones for detecting the same effect — Experiment Assignment Tracking, Metric Selection for Tests.

Excluding, without breaking the randomisation

The rule that makes exclusions safe:

Exclude on properties that cannot have been affected by the treatment. Bot user-agent, internal IP, employee account flag, assignment timestamp — all fixed independently of which variant someone got. Excluding on those is fine and doesn’t bias anything.

Never exclude on behaviour. “Remove users who bounced immediately” or “exclude sessions under 5 seconds” filters on an outcome the treatment may have caused, which breaks the comparability randomisation bought you and can manufacture an effect from nothing — Selection Bias.

✓  exclude: user_agent matches known bot list
✓  exclude: IP in office range
✓  exclude: account flagged internal
✓  exclude: assigned before the QA sign-off timestamp

✗  exclude: session under 5 seconds        ← treatment may cause short sessions
✗  exclude: users who saw an error         ← treatment may cause errors
✗  exclude: "outlier" order values          ← unless capped symmetrically and
                                              declared in advance

Apply exclusions identically to both arms, and check the counts after excluding — if the exclusion removed 8% of control and 12% of variant, the exclusion itself is treatment-dependent and you’ve just created the problem you were fixing.

Catching it

  • Sample Ratio Mismatch check first, always. The single most informative test of whether the sample is trustworthy, and it catches most pollution that matters
  • An A test before trusting a new platform — it surfaces duplicate assignment and filtering asymmetries with no real change to confuse things
  • Compare assigned counts to exposed counts. A large gap is the dilution problem above
  • Watch the conversion rate of the whole test population against the site baseline. Far below it usually means bots; far above usually means you’ve excluded the wrong people
  • Segment by bot-likelihood as a diagnostic, not as a filter — if excluding suspected bots changes the result materially, the result is fragile either way

Where it interacts

  • Bot and Internal Traffic — the analytics-side treatment of the same population, and the two systems should agree on the definition
  • Randomisation Unit — duplicate assignment is a consequence of choosing a unit you can’t identify reliably
  • Experiment QA — most pollution is preventable before launch and undetectable afterwards
  • Inconclusive Results — pollution is one of the standard explanations for a null, and it’s the one nobody checks