Tags: experimentation concept
Personalisation Tests
Date: 2026-08-17
Testing a targeting rule rather than an experience. The unit under test is “show X to segment S”, which means the segment definition is part of the treatment — and the standard mistake is comparing targeted users to untargeted ones, which compares two different populations and measures nothing.
A personalisation test tests a targeting rule — a given experience shown to a given segment — against not applying that rule.
The mistake that invalidates most of them
✗ WRONG
personalised experience shown to "high-intent users"
compared against everyone else
high-intent users convert at 8%
everyone else converts at 2%
"personalisation delivered a 4× lift" ← measures the segment,
not the personalisation
✓ RIGHT
randomise WITHIN the segment
high-intent users
├─ 50% personalised experience 8.4%
└─ 50% default experience 8.0% ← the control that matters
effect of personalisation on this segment: +5% relative
The control must be the same people, not the other people. Everything else in this note follows from that.
The design
The clean structure is a two-level randomisation, because there are two separate questions.
all traffic
│
randomise into the test
│
┌────────────────┴────────────────┐
│ │
TREATMENT arm CONTROL arm
targeting rule ACTIVE everyone gets default
│ │
┌────┴────┐ ┌────┴────┐
│ │ │ │
in segment not in in segment not in
→ variant → default → default → default
question 1: does the rule help the people it targets?
in-segment treatment vs in-segment control
question 2: does the rule help overall?
whole treatment arm vs whole control arm
← this is the one the business is buying
Keep the control arm’s segment membership computed but unused. You need to know who would have been targeted in order to answer question 1, so the segmentation logic must run in both arms and only act in one.
Question 2 is the one that gets skipped
A targeting rule can help its segment and be worthless overall, and this is common rather than exotic.
segment = "viewed 3+ products", 12% of traffic
personalised experience for them: +6% relative conversion
overall effect = 0.12 × 6% = +0.72% relative
In plain terms: a 6% win on 12% of traffic is a 0.72% win on the site — and 0.72% may well be inside the noise of the overall test. That’s the number to power the test on if the decision is “do we ship this rule”, because shipping affects the whole site’s numbers and its whole maintenance burden.
It gets worse when the rule can harm the untargeted:
in-segment +6% on 12% of traffic = +0.72%
out-of-segment −1% on 88% of traffic = −0.88% ← the default page
lost a slot to
the personalisation block
net −0.16%
Always measure the out-of-segment arm. Personalisation frequently costs the majority to serve the minority, and only the two-level design shows it.
Why the power is worse than you expect
Three penalties compound:
- The segment is a fraction of traffic. A 12% segment needs the whole test to run roughly 8× longer to get the same in-segment sample
- The overall effect is diluted by the segment share, as above, so detecting it needs a much smaller MDE than the in-segment effect suggests — Minimum Detectable Effect
- Multiple segments means multiple comparisons. Five rules tested at once is five tests, and the correction applies — The Multiple Comparisons Problem
segment share in-segment effect needed for a +1% overall effect
50% +2%
20% +5%
10% +10%
5% +20% ← implausible. the rule cannot pay for itself
Below roughly 10% of traffic, a targeting rule has to produce an enormous in-segment effect to be worth the machinery. That table is the most useful thing to have in front of you when someone proposes a fifteenth segment.
Segment definitions are part of the treatment
- Define the segment before launch, precisely, in the pre-registration. Adjusting it after seeing results is post-hoc segmentation wearing a product hat, and it manufactures findings — Pre-Registration
- Segments computed from live behaviour drift. “High intent” recomputed weekly means the population changes mid-test, so the thing you tested is not the thing you ship — Metric Drift
- Membership must be determinable at exposure time. A rule needing data that arrives later can’t be applied in real time, and testing it offline answers a different question
- A user moving between segments mid-test needs a documented rule — usually “first assignment sticks”, or the analysis becomes intractable
What a genuine win looks like
If in-segment and out-of-segment effects differ substantially and the difference has a mechanism, that’s a real heterogeneous treatment effect — the strongest justification for personalisation and rarer than the industry implies. Most measured segment differences are noise.
The honest comparison for any personalisation programme is against the simplest alternative: showing the best-performing single experience to everyone. Personalisation must beat that, not beat the old default — and it has to beat it by enough to cover the ongoing cost of maintaining rules, segments and the content variants they consume.
Where it interacts
- Heterogeneous Treatment Effects — the phenomenon that justifies personalisation, and the standard for proving it exists
- Segmentation (test results) — the difference is timing: a segment declared before launch is a design, the same segment found afterwards is noise
- Multi-Armed Bandits — contextual bandits are personalisation with automatic allocation, and inherit both sets of problems
- Interaction Effects — targeting rules layered over other tests are a systematic source of real interactions, because targeting isn’t randomised