Tags: statistics experimentation concept
Sample Ratio Mismatch
Date: 2026-08-16
The split you got doesn’t match the split you asked for. It means assignment or measurement is broken, which means the two groups aren’t comparable, which means the result is unreadable — however good it looks.
What it is
Sample ratio mismatch (SRM) is a statistically significant deviation between the observed allocation and the intended one. You asked for 50/50; you got 52/48 on large numbers, and that gap is too large to be chance.
It’s the highest-value check in experimentation because it’s cheap, automatic, and it invalidates everything downstream.
The check
A chi-squared goodness-of-fit test against the expected split. Worked, on a test that looks like a clear win:
observed expected (50/50)
control 51,240 52,634
variant 54,028 52,634
─────── ───────
total 105,268 105,268
χ² = Σ (O − E)² / E
= (51,240 − 52,634)² / 52,634 + (54,028 − 52,634)² / 52,634
= 1,943,236 / 52,634 + 1,943,236 / 52,634
= 36.92 + 36.92
= 73.8
With 1 degree of freedom, χ² = 73.8 gives p < 0.0001.
A 51,240 / 54,028 split looks like nothing — a 2.6 percentage point difference in allocation, the sort of thing that reads as rounding. It is not rounding. At this sample size, a fair coin produces a gap that large essentially never.
In plain terms: if assignment were working, you would not see this many more people in one group. Something is systematically sending users one way or removing them from the other — and whatever that something is, it also decides who ends up in which group, so the groups now differ by more than the change you’re testing.
Threshold: treat p < 0.001 as an SRM. A looser threshold produces false alarms across a busy programme; this is one place where being conservative is right.
Why it invalidates the whole result
The missing users aren’t missing at random. Control is roughly 1,400 short of where it should be, and those users went for a reason — a redirect that failed on slow connections, a script that errored on older browsers, a cache that served one variant preferentially. Those users differ from the ones who remained, and they converted differently.
So the comparison is no longer variant-versus-control. It’s variant-versus-a-filtered-subset-of-control, and the filter correlates with behaviour.
You cannot fix it by re-weighting. Scaling control back up to 50% assumes the missing users resembled the present ones, which is exactly the assumption the SRM tells you is false. The only response is to find the cause, fix it, and rerun.
Causes, roughly by frequency
- Redirect-based tests — the variant is a separate URL and some users never arrive. The classic cause, and the reason Split URL Tests are riskier than they look
- Asymmetric tracking loss — the variant loads an extra script that ad blockers catch, so its users are under-recorded
- Caching — a CDN serves one variant to users who should have got the other, or caches the assignment itself
- Bot filtering applied after assignment — bots split evenly, get removed unevenly because they behave differently in each arm — Bot and Internal Traffic
- Assignment bugs — a hash that isn’t uniform, a fallback that defaults to control on error, a race between assignment and page render
- Late-loading variant code — users who leave before the variant applies get counted as control
- Trigger conditions differing between arms — the exposure event fires on different criteria in each variant
In practice
- Automate it. Check daily, alert on breach, and make it the first line of every results readout — Guide - Statistics for CRO
- Check it during the test, not after. This is the one metric you monitor continuously without it counting as Peeking, because you’re not looking at the outcome
- Check unequal splits too. A 90/10 ramp has an expected ratio like any other
- Check within segments if the overall ratio is clean. A device-specific redirect failure can cancel out in the total
- Record SRM failures in the archive. They’re the most informative failures you’ll have — each one is a real bug in the delivery path that was also affecting users outside the test
Filed in Statistics rather than Experimentation because it’s a goodness-of-fit test, and because the same check applies anywhere allocation is supposed to be random.