Tags: statistics concept
Simpson’s Paradox
Date: 2026-08-16
Every segment improves and the total gets worse. Not a trick or an error — both readings are arithmetically correct, and which one is right depends on a causal question the numbers can’t answer.
What it is
Simpson’s paradox is when a trend that holds inside every subgroup reverses or disappears once the subgroups are combined.
It happens whenever two things are compared as rates, the subgroups have different baseline rates, and the mix of subgroups differs between the two things being compared. The aggregate rate is an average weighted by group size, so changing the weights moves it independently of anything happening inside the groups.
Worked
A test where the variant wins on mobile, wins on desktop, and loses overall.
MOBILE DESKTOP
visitors orders CR visitors orders CR
control 10,000 200 2.00% control 40,000 1,600 4.00%
variant 40,000 840 2.10% variant 10,000 410 4.10%
COMBINED
visitors orders CR
control 50,000 1,800 3.60%
variant 50,000 1,250 2.50% ← variant loses by 31%
Check the arithmetic: variant beats control by 5% relative on mobile (2.10 vs 2.00) and by 2.5% relative on desktop (4.10 vs 4.00). Combined, it converts at 2.50% against 3.60%.
Every number is correct. Nothing has been miscounted.
The mechanism
The segments have very different baseline rates — desktop converts at twice mobile — and the mix differs between arms. Control’s traffic is 80% desktop; the variant’s is 80% mobile. The variant is loaded with the low-converting segment, so its total is dragged down regardless of it performing better within each one.
control mix: 20% mobile / 80% desktop → weighted towards the good segment
variant mix: 80% mobile / 20% desktop → weighted towards the poor segment
In plain terms: the aggregate number is an average weighted by group size. Change the weights and the average moves, even when nothing inside the groups moves at all. Two things are being compared — performance within groups, and which groups people are in — and the total confuses them.
Which reading is right
This is the part people skip. The paradox isn’t resolved by statistics; it’s resolved by asking what caused the mix difference.
| If the mix difference is… | Then | Because |
|---|---|---|
| A bug — assignment isn’t random by device | Segments are right, total is meaningless | This is Sample Ratio Mismatch, and the test is invalid. Fix and rerun |
| Caused by the variant — it drove desktop users away before exposure | Total is right, segments mislead | Driving away good traffic is a real effect of your change, and it’s a loss |
| Pre-existing — the segments are just different populations | Segments are right; report them separately | The total answers a question nobody asked |
In A/B testing the first is overwhelmingly the most common, which is why the practical response to a Simpson’s-shaped result is check the SRM before anything else.
Where it appears outside tests
- Channel performance. Every channel’s conversion rate improves while site-wide conversion falls, because a cheap high-volume channel grew its share
- Cohort comparisons. Every acquisition cohort retains better than the last, while overall retention declines, because recent cohorts are larger and younger — Cohort Analysis
- Year-on-year metrics. Every product category grew its margin while blended margin fell, because the low-margin category grew fastest
- Before/after a redesign. Traffic mix shifted at the same time, so nothing is comparable — see Metric Drift, and Metric Decomposition to split the change into mix and rate
The common structure: a rate compared across two periods or groups whose composition changed. Whenever you see one, the mix is the first thing to check.
Defences
- Look at segments as a matter of routine, but pre-register which ones — post-hoc slicing has its own problem (Segmentation (test results), The Multiple Comparisons Problem)
- Check the mix explicitly. Compare segment proportions between groups before comparing rates
- Randomise properly. Correct randomisation makes mix differences vanish at scale, which is exactly why a mix difference is a red flag rather than a curiosity — Why Randomisation Works
- Stratify when the mix is genuinely unequal. Compute the effect within each segment and combine with fixed weights
- Report absolute counts alongside rates. The paradox is much harder to hide when the group sizes are on the page