Tags: commerce experimentation concept
Incrementality Testing
Date: 2026-08-16
Withhold the spend from a random group and measure the difference. It’s the only method that answers what a channel caused rather than what it was near — and it consistently finds that the most-credited channels are the least incremental.
What it is
Incrementality testing measures the causal effect of marketing spend by randomly withholding it from a comparable group and comparing outcomes.
ATTRIBUTION INCREMENTALITY
"how many conversions "how many conversions
had contact with this would NOT have happened
channel?" without this channel?"
answers: correlation answers: causation
Every attribution model — last click, multi-touch, data-driven — answers the first question. None of them answers the second, because the counterfactual isn’t in observational data. See Attribution Models, Multi-Touch Attribution.
The mechanism
It’s a controlled experiment where the treatment is exposure to advertising:
randomly split the eligible population
treated sees the ads conversion rate 4.20%
holdout ads suppressed conversion rate 3.90%
──────
incremental lift 0.30pp
On 250,000 in each group, that’s 750 incremental conversions. Compare against what the platform claimed:
platform-reported conversions 2,400
genuinely incremental 750
─────
incrementality 31%
Two-thirds of the reported conversions would have happened anyway. That’s a typical shape of finding, not a worst case — and it changes true CAC by a factor of three.
Why bottom-funnel channels are worst
The mechanism is structural, not vendor misbehaviour.
Retargeting and branded search both work by selecting people who have already shown intent. Someone who visited a product page yesterday is likely to buy whether or not you show them an ad. Someone searching your brand name was already coming.
So those channels look excellent in every attribution model — they’re present on converting paths by construction — and they’re the least incremental. The correlation is real and the causation isn’t.
Upper-funnel channels have the opposite profile: genuinely incremental, and invisible to last-click attribution.
The designs
| Design | Splits by | Suits |
|---|---|---|
| Audience holdout | Users, inside the ad platform | Where the platform supports it |
| Geo holdout | Regions | Channels with no user-level control |
| Ghost ads / PSA | Users, showing a placebo | Cleanest, needs platform support |
| Spend on/off over time | Time periods | Weakest — confounded with everything seasonal |
The last is common and barely valid: turning spend off for a fortnight compares two different fortnights, confounded with everything that changed. Use it only when nothing else is available, and read it as directional.
The costs
Real, and worth stating plainly:
- You deliberately lose sales. The holdout group isn’t advertised to, and some of them don’t buy. That’s the price of knowing
- Power. Incremental effects are small relative to baseline conversion, so these tests need large populations or long periods — often larger than an on-site A/B test — Minimum Detectable Effect
- Geo designs have tiny effective sample sizes. The unit is the region, not the user — twenty regions means n = 20, not n = 500,000
- Contamination. Holdout users may still see organic mentions, or the brand elsewhere
When it’s worth it
Where the spend is large enough that being wrong by 3× matters. For a £20,000 annual channel, run it on judgement. For £500,000, the test pays for itself many times over.
And run it before a budget increase, not after. The most common use is defensive — a channel looks efficient, someone proposes doubling it, and the holdout establishes whether the efficiency is real before the money moves.
Feeds Marketing Mix Modelling for the portfolio view, and it’s the ground truth that Modelled Conversions and Walled Garden Reporting can only approximate.