Tags: statistics commerce concept

Difference-in-Differences

Date: 2026-08-17


Using an untreated group’s change over the same period as the counterfactual for the treated group’s change. It rescues causal inference where randomisation was impossible — regional rollouts, pricing changes, TV campaigns — and rests entirely on one assumption you can support but never prove.


Difference-in-differences (DiD) estimates an effect by subtracting a comparison group’s before-to-after change from the treated group’s before-to-after change.

The arithmetic

Two groups, two periods, four numbers.

a delivery-threshold change rolled out to the North region only

                    BEFORE    AFTER    change
North (treated)      4.0%      4.9%    +0.9pp
South (control)      3.6%      4.1%    +0.5pp
                                       ───────
DiD estimate  =  0.9 − 0.5  =  +0.4pp

Read it as: the North rose 0.9pp, but 0.5pp of that would have happened anyway — because the South, which had no change, rose 0.5pp over the same period. The intervention is credited with the remaining 0.4pp.

Compare what the two naive readings would have said:

before/after on the treated group alone      +0.9pp   ← 2.25× overstated
treated vs control, after period only        +0.8pp   ← also wrong; the
                                                        regions differed
                                                        by 0.4pp to begin with
difference-in-differences                    +0.4pp

The first differencing removes fixed differences between the groups; the second removes changes common to both. That’s the whole design, and it’s why the method works even when the groups were never comparable in level.

The assumption: absent the intervention, the two groups’ outcomes would have moved in parallel. Not that they’d be equal — that they’d change by the same amount.

        conversion
   5% ┤                              ╭─ North actual (4.9%)
      │                          ╭───╯
      │                      ╭───╯ ← the +0.4pp attributed
   4% ┤      ╭───────────╮───╯···· North counterfactual (4.5%)
      │  ╭───╯       ╭───╯
      │──╯       ╭───╯                South actual (4.1%)
   3% ┤──────────╯
      └──┬────┬────┬────┬────┬────┬──
        −4   −3   −2   −1    0   +1   period, intervention at 0

You cannot test the assumption for the period that matters — the counterfactual is unobservable by definition. What you can do:

  • Plot several pre-periods. If the lines moved in parallel for six months before, parallel trends is credible. If they were diverging already, the estimate absorbs that divergence and is wrong
  • Run a placebo test. Apply the same DiD to a pre-intervention period where nothing happened. It should return approximately zero. If it returns a large “effect”, the design is broken
  • Test other outcomes the intervention shouldn’t have affected. A significant DiD on an unrelated metric means something else differed between the regions

What breaks it

  • Non-parallel pre-trends. The single most common failure, and the reason to always plot before computing
  • Something else changed in one group. A competitor opened in the South, a distribution centre closed in the North, one region had different weather. DiD attributes every group-specific concurrent change to your intervention
  • Composition shifts. If the treatment changed who visits — a campaign that brought new traffic to the North — the groups’ underlying mixes diverge and the comparison is no longer like-for-like
  • Spillover. If the control group is affected by the treatment, it’s not a control. Neighbouring regions, shared media, customers who move between them
  • Choosing the control after seeing results. Trying six candidate control regions and keeping the one that gives a clean answer is The Garden of Forking Paths — pick the control and the pre-period before looking

Getting the uncertainty right

The four-number table gives a point estimate with no interval, and the interval is where DiD analyses most often mislead.

Standard errors must account for correlation within groups over time. Weekly observations from the same region aren’t independent — a good week in the North is followed by another good week. Treating 52 weekly observations as 52 independent points understates the standard error badly, often by a factor of two or more.

In plain terms: with two regions you effectively have two observations, not 104. This is why credible DiD work uses many treated and control units — twenty regions, not one — and why a two-region DiD should be reported as suggestive rather than as a measurement.

Where it earns its place

  • Geographic rollouts — pricing, delivery terms, a new store format. The natural application — Geo Holdout Tests
  • Media that can’t be split by user — TV, radio, out-of-home, where geography is the only available randomisation unit
  • Policy or platform changes affecting some customers and not others for reasons unrelated to their behaviour
  • Retrospective evaluation of something already rolled out, where a proper test wasn’t run — the common real case, and the one to be most cautious with

Where a geo holdout is available, randomise which regions are treated. That converts DiD from an observational method resting on parallel trends into a genuine experiment where the assumption holds by construction — the same upgrade randomisation always provides, and it’s usually cheap to arrange in advance and impossible to arrange afterwards.

Where it interacts

  • Counterfactuals — DiD is one construction of the missing comparison, and parallel trends is its price
  • Geo Holdout Tests — the commerce application, and the randomised version of this design
  • Marketing Mix Modelling — the alternative when there’s no clean untreated group, with weaker causal standing
  • Natural Experiments — DiD is often the analysis method applied to one