Tags: experimentation concept
Why Randomisation Works
Date: 2026-08-16
It balances the variables you never thought to measure. Every other method of forming comparison groups only balances the ones you listed — and the confounder that ruins your result is always one you didn’t list.
What it is
Randomisation is assigning each unit to a group by a process nobody and nothing can influence, so group membership is statistically independent of every characteristic the unit has.
That independence is the whole mechanism. It isn’t about fairness or about avoiding bias in the colloquial sense — it’s a mathematical property that makes the two groups exchangeable in expectation.
The property
Because assignment is independent of everything, every attribute distributes evenly across arms as sample size grows — measured, unmeasured, and unimaginable.
control variant
mobile share 52.1% 51.8% ← you'd have thought to balance this
new visitors 68.4% 68.9% ← and this
average past order value £47.20 £47.90 ← maybe this
...
intent to buy today ? ? ← balanced too, and unmeasurable
mood ? ? ← balanced too
already saw the ad ? ? ← balanced too
The bottom three are the point. You cannot measure them, cannot match on them, and cannot adjust for them — and randomisation handles them anyway, at no cost and without you knowing what they are.
In plain terms: matching two groups by hand means matching on the things you can think of. Randomising means the things you can’t think of get split evenly too, which is the only reason the comparison is trustworthy.
Why the alternatives fail
| Method | Balances | Misses |
|---|---|---|
| Before / after | Nothing | Time, campaigns, seasonality, mix, everything |
| Users who chose the feature | Nothing | Self-selection — they were already different |
| Match on device and location | Device and location | Every other variable, including intent |
| Alternate every other visitor | Roughly, arrival order | Anything correlated with sequence — and it’s predictable, so it’s game-able |
| Split by region | Nothing reliably | Regions genuinely differ; n = number of regions, not users |
The pattern: each balances the listed variables and leaves the unlisted ones to chance in a way that isn’t random. See Confounding Variables and Selection Bias.
”In expectation” is doing work
Randomisation balances groups on average across repeated experiments, not perfectly in any one. A single test can land with more mobile users in one arm by luck — that’s not a broken randomiser, it’s ordinary variation, and it’s exactly what the confidence interval accounts for.
Two practical consequences:
- Small samples randomise badly. With 200 users per arm, meaningful imbalance is common. Randomisation’s guarantee is asymptotic
- Checking balance post hoc is of limited use. Finding an imbalanced covariate and “correcting” for it reintroduces the analyst’s judgement that randomisation existed to remove. Check it, note it, but the pre-registered analysis stands — see Pre-Registration
The exception worth knowing: stratified assignment randomises separately within groups (device, new/returning), guaranteeing balance on those while keeping randomisation for everything else. Strictly better where the strata are known and few — see Variance Reduction.
How it breaks in practice
Randomisation is fragile in implementation even when it’s sound in principle.
- A hash that isn’t uniform — a poor bucketing function correlates with something in the identifier, like signup date. See Assignment and Bucketing
- Assignment on something non-random — session ID that resets, a cookie that clears unevenly, a user ID that encodes cohort
- Fallback to control on error — users whose assignment call fails land in control, and they fail for reasons (slow connections, old browsers) that predict conversion. This shows up as Sample Ratio Mismatch
- Re-randomising returning users — assignment must be sticky, or a user contributes to both arms and the arms are no longer independent
- Filtering after assignment — removing bots or internal traffic post-hoc is fine only if the filter is blind to arm
The single check that catches most of these is SRM, which is why it’s the first thing to look at in any result — Reading a Test Result.