Tags: experimentation statistics concept
Heterogeneous Treatment Effects
Date: 2026-08-17
A treatment genuinely affecting different groups by different amounts, rather than the average effect applying uniformly. It’s real, it’s the justification for personalisation, and it’s almost impossible to distinguish from noise using the same test that measured the average — which is why nearly every claimed instance is a false discovery.
A heterogeneous treatment effect (HTE) is a treatment effect that genuinely differs in size or direction between groups of users, so the average effect describes none of them exactly.
Real, versus what it usually is
GENUINE HTE POST-HOC NOISE
declared before launch found by slicing afterwards
one or two pre-specified dimensions fourteen dimensions, checked
mechanism stated in advance mechanism invented to fit
powered for the subgroup subgroup has 8% of the sample
replicates in a second test never re-tested
the INTERACTION is significant, only the subgroup's own effect
not just the subgroup's effect is "significant"
The right-hand column is what “it worked really well on mobile” almost always is — Segmentation (test results).
The test people don’t run
The critical technical point. Finding that a subgroup’s effect is significant while another’s isn’t does not establish that they differ.
mobile +6.2% 95% CI +1.1% to +11.5% ← "significant!"
desktop +2.1% 95% CI −2.4% to +6.8% ← "not significant"
conclusion drawn: "it works on mobile, not desktop"
conclusion warranted: none
the difference between them: +4.1pp
95% CI on that difference: −2.6pp to +10.8pp ← includes zero
→ no evidence they differ
You must test the interaction directly — the difference of differences — not compare two significance verdicts. Two intervals that overlap substantially cannot support a claim that the effects differ, and one being on either side of an arbitrary line says nothing.
In plain terms: “significant here, not there” is a statement about how much data each group had, not about whether the treatment behaves differently. Bigger groups produce narrower intervals, so the larger segment is more likely to show significance for the same true effect.
And the interaction test is expensive: detecting a difference of differences reliably takes roughly four times the sample of detecting a main effect of the same magnitude — Multivariate Tests, Interaction Effects.
Why the false discoveries are so persuasive
With ten segmenting dimensions at two levels each, there are twenty subgroups. At 95% confidence:
P(at least one false positive) = 1 − 0.95²⁰ = 64%
Nearly two tests in three will produce a subgroup that “worked” by chance — The Multiple Comparisons Problem.
Then three things make it stick. The story is always available: any subgroup finding can be explained after the fact, and mobile-versus-desktop has a plausible mechanism ready for either direction. The Winner’s Curse applies doubly — a subgroup effect selected for being large is inflated by more than a main effect would be, because the subgroup is smaller and noisier. And nobody re-tests it, because it already has a number attached.
Establishing one properly
- Declare the dimension before launch, with a mechanism — “delivery-date messaging should matter more to first-time buyers, because repeat buyers already know our delivery times”. One or two dimensions, not a menu
- Power for the subgroup, not the whole test. If first-time buyers are 30% of traffic, the test needs roughly 3× the sample to detect the same effect within them — and more again for the interaction
- Test the interaction term, and report the confidence interval on the difference
- Replicate. A subgroup finding that survives a second, pre-registered test is real. One that doesn’t was noise, and this is the only decisive check available
- Prefer dimensions fixed before assignment — device, new versus returning, acquisition channel, region. Anything measured after exposure may have been caused by the treatment, which is a different problem entirely — Selection Bias
The dimensions genuinely worth pre-specifying
Where real heterogeneity is most often found in commerce, and worth a mechanism when you propose one:
- Device — the most commonly real, because the experience genuinely differs. Also the most commonly claimed falsely
- New versus returning — different information, different habits, different priors about you
- Traffic source — paid-search intent differs from social browsing
- Purchase history / value tier — where the change touches price, delivery or loyalty
- Geography, where delivery, tax or regulation differ materially
What to do with a real one
Even when established, shipping a segment-specific experience isn’t automatic:
in-segment effect +6% on 25% of traffic
overall effect +1.5%
maintenance cost two experiences, forever, in every future change
versus shipping to everyone at the average effect: +1.8%
Sometimes the average effect is better than the personalised one once complexity is priced in. Uniform rollout is the default; heterogeneity has to beat it, not merely exist — Personalisation Tests.
Where it interacts
- Segmentation (test results) — the same slicing done after the fact, and why it produces noise
- Personalisation Tests — the design that acts on heterogeneity, and the correct way to verify it prospectively
- Winner’s Curse — why subgroup effects are systematically overstated even when real
- Simpson’s Paradox — the related failure where subgroup and aggregate effects point in opposite directions for compositional reasons