Tags: commerce experimentation concept
Geo Holdout Tests
Date: 2026-08-16
Withhold spend by region and compare. It’s the practical way to measure incrementality when you can’t randomise users — and its weakness is that your sample size is the number of regions, not the number of customers.
What it is
A geo holdout test randomises geographic areas into treated and control groups, runs spend in one and not the other, and compares outcomes.
UK regions, randomly assigned
treated spend continues revenue per capita £4.20
holdout spend suppressed revenue per capita £3.85
──────
incremental £0.35 per head
Used where user-level randomisation isn’t possible: TV, radio, out-of-home, and any ad platform that won’t hold out an audience.
The sample size trap
The most important property, and the one that catches people:
Your n is the number of regions, not the number of people in them.
20 regions, 2.5 million people
feels like n = 2,500,000
actually n = 20
Twenty observations is a very small experiment. Two regions behaving unusually — a local competitor opening, a weather event, a football result — can dominate the result, and there’s no amount of population inside a region that fixes it.
This is why geo tests need either many regions, long durations, or both, and why a test across four regions is essentially uninterpretable — Sampling Error.
Designing one properly
Match before you randomise. Regions differ enormously in baseline. Pairing similar regions and randomising within pairs removes most of that variance:
pair by pre-period revenue per capita, then randomise within pair
Manchester ↔ Leeds one treated, one held out
Bristol ↔ Nottingham
Glasgow ↔ Newcastle
This is stratification, and it’s close to essential here given how few units you have — Variance Reduction.
Use a pre-period. Measure both groups before the test starts. The comparison is then the change in the gap rather than the gap itself, which absorbs persistent regional differences — that’s Difference-in-Differences.
Run long enough to cross a full cycle, including at least one payday and no single unrepresentative event.
Watch for spillover. Regions aren’t sealed — people travel, national media reaches everywhere, and online word of mouth ignores geography. Spillover dilutes the measured effect, so a geo test tends to understate incrementality.
Analysing it
- Per-capita metrics, not totals, or region size dominates
- The unit of analysis is the region. Compute per-region outcomes, then compare across regions. Pooling all customers treats it as a user-level test and produces spuriously narrow intervals — the same error as Ratio Metrics with a varying denominator
- Difference-in-differences rather than a raw comparison
- Expect wide intervals. With n = 20 the uncertainty is genuinely large, and reporting a point estimate without it is the main way these get over-read — Communicating Uncertainty
Where it’s weak
- Low power. Often you can only detect large effects. A channel with a genuine but modest incremental effect may be indistinguishable from zero
- Confounding events. One region has a store closure, a local news story, a supply issue. With twenty units, one event matters
- Regional heterogeneity. The effect in cities may differ from rural areas, and a blended answer describes neither
- You lose sales in the holdout, deliberately
When to use it
- Offline and broadcast channels, where there’s no other option
- Platforms that won’t hold out an audience
- Validating a platform’s own incrementality claim with a method you control
- Before a major budget change, where being wrong by 3× is expensive
For channels where user-level holdout is available inside the platform, prefer that — the sample size is vastly better. Geo is the fallback, not the first choice — Incrementality Testing.