Tags: statistics concept

Simpson’s Paradox

Date: 2026-08-16


Every segment improves and the total gets worse. Not a trick or an error — both readings are arithmetically correct, and which one is right depends on a causal question the numbers can’t answer.


What it is

Simpson’s paradox is when a trend that holds inside every subgroup reverses or disappears once the subgroups are combined.

It happens whenever two things are compared as rates, the subgroups have different baseline rates, and the mix of subgroups differs between the two things being compared. The aggregate rate is an average weighted by group size, so changing the weights moves it independently of anything happening inside the groups.

Worked

A test where the variant wins on mobile, wins on desktop, and loses overall.

MOBILE                          DESKTOP
        visitors  orders  CR            visitors  orders  CR
control   10,000     200  2.00%  control   40,000   1,600  4.00%
variant   40,000     840  2.10%  variant   10,000     410  4.10%

COMBINED
        visitors  orders     CR
control   50,000   1,800   3.60%
variant   50,000   1,250   2.50%   ← variant loses by 31%

Check the arithmetic: variant beats control by 5% relative on mobile (2.10 vs 2.00) and by 2.5% relative on desktop (4.10 vs 4.00). Combined, it converts at 2.50% against 3.60%.

Every number is correct. Nothing has been miscounted.

The mechanism

The segments have very different baseline rates — desktop converts at twice mobile — and the mix differs between arms. Control’s traffic is 80% desktop; the variant’s is 80% mobile. The variant is loaded with the low-converting segment, so its total is dragged down regardless of it performing better within each one.

control mix:   20% mobile / 80% desktop  →  weighted towards the good segment
variant mix:   80% mobile / 20% desktop  →  weighted towards the poor segment

In plain terms: the aggregate number is an average weighted by group size. Change the weights and the average moves, even when nothing inside the groups moves at all. Two things are being compared — performance within groups, and which groups people are in — and the total confuses them.

Which reading is right

This is the part people skip. The paradox isn’t resolved by statistics; it’s resolved by asking what caused the mix difference.

If the mix difference is…ThenBecause
A bug — assignment isn’t random by deviceSegments are right, total is meaninglessThis is Sample Ratio Mismatch, and the test is invalid. Fix and rerun
Caused by the variant — it drove desktop users away before exposureTotal is right, segments misleadDriving away good traffic is a real effect of your change, and it’s a loss
Pre-existing — the segments are just different populationsSegments are right; report them separatelyThe total answers a question nobody asked

In A/B testing the first is overwhelmingly the most common, which is why the practical response to a Simpson’s-shaped result is check the SRM before anything else.

Where it appears outside tests

  • Channel performance. Every channel’s conversion rate improves while site-wide conversion falls, because a cheap high-volume channel grew its share
  • Cohort comparisons. Every acquisition cohort retains better than the last, while overall retention declines, because recent cohorts are larger and younger — Cohort Analysis
  • Year-on-year metrics. Every product category grew its margin while blended margin fell, because the low-margin category grew fastest
  • Before/after a redesign. Traffic mix shifted at the same time, so nothing is comparable — see Metric Drift, and Metric Decomposition to split the change into mix and rate

The common structure: a rate compared across two periods or groups whose composition changed. Whenever you see one, the mix is the first thing to check.

Defences

  • Look at segments as a matter of routine, but pre-register which ones — post-hoc slicing has its own problem (Segmentation (test results), The Multiple Comparisons Problem)
  • Check the mix explicitly. Compare segment proportions between groups before comparing rates
  • Randomise properly. Correct randomisation makes mix differences vanish at scale, which is exactly why a mix difference is a red flag rather than a curiosity — Why Randomisation Works
  • Stratify when the mix is genuinely unequal. Compute the effect within each segment and combine with fixed weights
  • Report absolute counts alongside rates. The paradox is much harder to hide when the group sizes are on the page