Tags: experimentation concept
Experiment QA
Date: 2026-08-16
The checks between “the variant is built” and “traffic is flowing”. Skipping them doesn’t produce a failed test — it produces a completed test with a confident wrong answer, discovered nine weeks later or never.
What it is
Experiment QA is verifying that the test measures what it claims to, before it starts. Distinct from testing the variant works — that’s ordinary QA — it’s about whether the experimental apparatus is sound.
The asymmetry that justifies the effort: a bug found before launch costs an hour. The same bug found after costs the entire test duration, and if it’s never found it costs a wrong belief that propagates into every decision downstream.
The four things that break
1. Assignment isn’t random or isn’t sticky
- Does the same user get the same variant across page loads, sessions, and devices where expected?
- Does the split land where configured? Check bucket distribution on a sample before launch, not just after — Assignment and Bucketing
- Does anything fall back to control on error? That loads control with slow connections and old browsers — Why Randomisation Works
- Does logging in reassign the user mid-visit?
2. Exposure isn’t tracked correctly
- Does the exposure event fire when the variant could first influence behaviour, not at session start? Dilution otherwise — Experiment Assignment Tracking
- Does it fire once per user, deduplicated?
- Does it fire identically in both arms? The single most common source of Sample Ratio Mismatch
- Can exposure be affected by the treatment? If the variant makes an element more prominent and exposure fires on interacting with it, the arms aren’t comparable
3. The variant doesn’t render everywhere it should
- Both arms, on the browsers and devices in your actual traffic — not just the two on the developer’s machine
- With ad blockers and privacy extensions active. If the variant needs a script that gets blocked, its users silently see control
- With consent denied. A test that only runs for consenting users is a test on a biased population — Consent Management
- On slow connections, where Flicker and Flash of Original Content appears
- Logged in and logged out; with an empty basket and a full one
4. The metrics don’t measure what the plan says
- Does the primary metric’s definition in the tool match the pre-registration — same denominator, same window, same timezone?
- Are guardrails configured and alerting, not just listed?
- Do conversions join to exposures? A backend purchase carrying only a user ID won’t join to a cookie-based exposure — Identity Stitching
The pre-launch sequence
Order matters — each step is cheaper than the one after it.
1 read the pre-registration back against the built test
2 force yourself into each variant (query param override) and walk the
full journey, including the failure paths
3 verify exposure and conversion events in a debug view, both arms
4 check bucket distribution on synthetic IDs — is the split uniform?
5 launch at 1% for a day
6 check SRM, exposure counts per arm, and guardrails before ramping
Step 5 is the one that gets skipped and shouldn’t. A day at 1% costs almost no traffic and catches the failures that only appear under real conditions — real devices, real blockers, real caching. Everything above it happens in an environment that isn’t production.
Ramp-day checks
At 1%, before opening it up:
| Check | Bad sign |
|---|---|
| Exposure counts per arm | Materially unequal — asymmetric tracking |
| Exposed vs assigned | Equal — exposure isn’t really being tracked |
| SRM | p < 0.001 — stop and diagnose |
| Guardrails | Any breach, at any significance |
| Error rate by arm | Variant higher — something’s throwing |
| Conversions joining | Zero conversions attributed — the join is broken |
Failure modes
- QA’d in staging only. Staging has no ad blockers, no CDN caching, no real device mix
- Only the variant checked. Control is a variant too, and control-side tracking regressions are common
- Checked on desktop Chrome. Where your traffic isn’t
- QA’d by the person who built it, who tests the paths they were thinking about
- No 1% ramp, so the first real-world signal arrives at full exposure
- Guardrails listed but not wired up, so nobody is watching the thing that was supposed to stop it
The procedure that assembles all of this: Guide - Running an Experiment.