Tags: commerce experimentation concept

Geo Holdout Tests

Date: 2026-08-16


Withhold spend by region and compare. It’s the practical way to measure incrementality when you can’t randomise users — and its weakness is that your sample size is the number of regions, not the number of customers.


What it is

A geo holdout test randomises geographic areas into treated and control groups, runs spend in one and not the other, and compares outcomes.

UK regions, randomly assigned

  treated   spend continues       revenue per capita  £4.20
  holdout   spend suppressed      revenue per capita  £3.85
                                                     ──────
  incremental                                        £0.35 per head

Used where user-level randomisation isn’t possible: TV, radio, out-of-home, and any ad platform that won’t hold out an audience.

The sample size trap

The most important property, and the one that catches people:

Your n is the number of regions, not the number of people in them.

20 regions, 2.5 million people

  feels like    n = 2,500,000
  actually      n = 20

Twenty observations is a very small experiment. Two regions behaving unusually — a local competitor opening, a weather event, a football result — can dominate the result, and there’s no amount of population inside a region that fixes it.

This is why geo tests need either many regions, long durations, or both, and why a test across four regions is essentially uninterpretable — Sampling Error.

Designing one properly

Match before you randomise. Regions differ enormously in baseline. Pairing similar regions and randomising within pairs removes most of that variance:

pair by pre-period revenue per capita, then randomise within pair

  Manchester ↔ Leeds        one treated, one held out
  Bristol    ↔ Nottingham
  Glasgow    ↔ Newcastle

This is stratification, and it’s close to essential here given how few units you have — Variance Reduction.

Use a pre-period. Measure both groups before the test starts. The comparison is then the change in the gap rather than the gap itself, which absorbs persistent regional differences — that’s Difference-in-Differences.

Run long enough to cross a full cycle, including at least one payday and no single unrepresentative event.

Watch for spillover. Regions aren’t sealed — people travel, national media reaches everywhere, and online word of mouth ignores geography. Spillover dilutes the measured effect, so a geo test tends to understate incrementality.

Analysing it

  • Per-capita metrics, not totals, or region size dominates
  • The unit of analysis is the region. Compute per-region outcomes, then compare across regions. Pooling all customers treats it as a user-level test and produces spuriously narrow intervals — the same error as Ratio Metrics with a varying denominator
  • Difference-in-differences rather than a raw comparison
  • Expect wide intervals. With n = 20 the uncertainty is genuinely large, and reporting a point estimate without it is the main way these get over-read — Communicating Uncertainty

Where it’s weak

  • Low power. Often you can only detect large effects. A channel with a genuine but modest incremental effect may be indistinguishable from zero
  • Confounding events. One region has a store closure, a local news story, a supply issue. With twenty units, one event matters
  • Regional heterogeneity. The effect in cities may differ from rural areas, and a blended answer describes neither
  • You lose sales in the holdout, deliberately

When to use it

  • Offline and broadcast channels, where there’s no other option
  • Platforms that won’t hold out an audience
  • Validating a platform’s own incrementality claim with a method you control
  • Before a major budget change, where being wrong by 3× is expensive

For channels where user-level holdout is available inside the platform, prefer that — the sample size is vastly better. Geo is the fallback, not the first choice — Incrementality Testing.