Tags: commerce experimentation concept

Incrementality Testing

Date: 2026-08-16


Withhold the spend from a random group and measure the difference. It’s the only method that answers what a channel caused rather than what it was near — and it consistently finds that the most-credited channels are the least incremental.


What it is

Incrementality testing measures the causal effect of marketing spend by randomly withholding it from a comparable group and comparing outcomes.

ATTRIBUTION                      INCREMENTALITY

"how many conversions            "how many conversions
 had contact with this            would NOT have happened
 channel?"                        without this channel?"

answers: correlation             answers: causation

Every attribution model — last click, multi-touch, data-driven — answers the first question. None of them answers the second, because the counterfactual isn’t in observational data. See Attribution Models, Multi-Touch Attribution.

The mechanism

It’s a controlled experiment where the treatment is exposure to advertising:

randomly split the eligible population

  treated   sees the ads          conversion rate  4.20%
  holdout   ads suppressed        conversion rate  3.90%
                                                  ──────
  incremental lift                                 0.30pp

On 250,000 in each group, that’s 750 incremental conversions. Compare against what the platform claimed:

platform-reported conversions    2,400
genuinely incremental              750
                                 ─────
incrementality                     31%

Two-thirds of the reported conversions would have happened anyway. That’s a typical shape of finding, not a worst case — and it changes true CAC by a factor of three.

Why bottom-funnel channels are worst

The mechanism is structural, not vendor misbehaviour.

Retargeting and branded search both work by selecting people who have already shown intent. Someone who visited a product page yesterday is likely to buy whether or not you show them an ad. Someone searching your brand name was already coming.

So those channels look excellent in every attribution model — they’re present on converting paths by construction — and they’re the least incremental. The correlation is real and the causation isn’t.

Upper-funnel channels have the opposite profile: genuinely incremental, and invisible to last-click attribution.

The designs

DesignSplits bySuits
Audience holdoutUsers, inside the ad platformWhere the platform supports it
Geo holdoutRegionsChannels with no user-level control
Ghost ads / PSAUsers, showing a placeboCleanest, needs platform support
Spend on/off over timeTime periodsWeakest — confounded with everything seasonal

The last is common and barely valid: turning spend off for a fortnight compares two different fortnights, confounded with everything that changed. Use it only when nothing else is available, and read it as directional.

The costs

Real, and worth stating plainly:

  • You deliberately lose sales. The holdout group isn’t advertised to, and some of them don’t buy. That’s the price of knowing
  • Power. Incremental effects are small relative to baseline conversion, so these tests need large populations or long periods — often larger than an on-site A/B test — Minimum Detectable Effect
  • Geo designs have tiny effective sample sizes. The unit is the region, not the user — twenty regions means n = 20, not n = 500,000
  • Contamination. Holdout users may still see organic mentions, or the brand elsewhere

When it’s worth it

Where the spend is large enough that being wrong by 3× matters. For a £20,000 annual channel, run it on judgement. For £500,000, the test pays for itself many times over.

And run it before a budget increase, not after. The most common use is defensive — a channel looks efficient, someone proposes doubling it, and the holdout establishes whether the efficiency is real before the money moves.

Feeds Marketing Mix Modelling for the portfolio view, and it’s the ground truth that Modelled Conversions and Walled Garden Reporting can only approximate.