Tags: experimentation statistics concept

Segmentation (test results)

Date: 2026-08-16


Slicing a finished test to find where it worked. Almost every finding from it is noise, because each slice is another comparison and the slices are chosen after seeing which ones look interesting.


What it is

Post-hoc segmentation is analysing a test result within subgroups that were not nominated before launch — mobile versus desktop, new versus returning, by channel, by basket value.

It is different from a pre-registered segment, which is part of the hypothesis and is a legitimate result. The distinction is entirely about when the segment was chosen, not what it contains.

The arithmetic

Six segmentation dimensions, each with two or three levels, is easily thirty comparisons:

device      mobile / desktop / tablet          3
visitor     new / returning                    2
channel     paid / organic / direct / email    4
basket      above / below AOV                  2
region      UK / rest of world                 2
logged in   yes / no                           2
                                             ───
                          3+2+4+2+2+2  =      15 slices, minimum

P(at least one significant by chance) = 1 − 0.95¹⁵ = 53.7%

More likely than not, one segment will look significant in a test where nothing happened. Cross two dimensions together — mobile and paid — and the count multiplies rather than adds. See The Multiple Comparisons Problem.

In plain terms: if you cut the data enough ways, one of the cuts will look like a discovery. That’s guaranteed by the threshold, not evidence against it.

Why it feels so convincing

The failure isn’t that people don’t know about multiplicity. It’s that a post-hoc segment always arrives with a story attached.

“It won on mobile and lost on desktop — of course, the desktop layout has more space so the change mattered less there.”

That explanation is plausible, generated after the fact, and would have been equally plausible in reverse. A mechanism invented to fit a finding isn’t evidence for it — see The Garden of Forking Paths. The test for whether you’d have predicted it: would you have written that sentence before launch? If not, treat it as a hypothesis.

Three further reasons segments mislead:

  • Each slice is underpowered. A test powered for 53,000 per arm, cut into mobile and desktop, gives you two underpowered tests. Effects that cross the line in a small slice are inflated — Winner’s Curse
  • Segments are correlated. Mobile, paid and new visitor are largely the same people, so “it won across three segments” is often one finding counted three times
  • Simpson’s Paradox means a segment can move opposite to the total for reasons of composition rather than behaviour

Doing it properly

Segmentation is genuinely valuable — the alternative isn’t refusing to look.

  • Pre-register the segments that matter, in the plan. Two or three, with a mechanism for why the effect should differ. Those are results
  • Always include new versus returning as a standing diagnostic, because it detects Novelty and Primacy Effects rather than exploring
  • Correct for the rest. If you’re going to slice fifteen ways, apply a false discovery rate correction — Bonferroni and False Discovery Rate
  • Treat unregistered findings as the next hypothesis, never as this test’s result. Write them into the Experiment Archive as candidates
  • Replicate before believing. A segment effect confirmed by a dedicated follow-up test targeting only that segment is real. Nothing else settles it
  • Prefer Heterogeneous Treatment Effects methods where you genuinely expect the effect to vary — they model variation across the whole population rather than testing arbitrary cuts

The rule of thumb

A segment result is a finding if you’d have bet on it beforehand, and a hypothesis if you wouldn’t.

The practical consequence for a readout: report the pre-registered segments as results, and anything else in a clearly-labelled “worth investigating” section that carries no decision. Making that separation visible on the page is most of the discipline, because it removes the ambiguity that lets a post-hoc slice quietly become the headline.