Tags: statistics concept

Propensity Score Matching

Date: 2026-08-17


Building a comparison group from observational data by pairing treated units with untreated ones that look similar. It’s genuinely useful and routinely oversold — it can only balance the variables you measured, so it fixes the confounding you already knew about and does nothing for the confounding that was the actual problem.


Propensity score matching estimates each unit’s probability of having been treated from its characteristics — its propensity score — and compares treated units with untreated ones of similar score.

The method

1  MODEL     estimate P(treated | characteristics) for every unit
             — usually Logistic Regression
             this probability is the PROPENSITY SCORE

2  MATCH     pair each treated unit with untreated unit(s) of
             similar score — nearest neighbour, calipers, or
             weight the whole sample by it

3  CHECK     verify the matched groups now look alike on every
             covariate    ← the step that gets skipped

4  COMPARE   difference in outcome between matched groups

The insight that makes it work: you don’t need to match on twenty variables individually. Units with the same probability of being treated have, on average, the same distribution of all the variables that went into computing it. One number stands in for the whole set.

Worked

Loyalty programme members appear to be worth far more:

                          members    non-members
                           24,000        180,000
avg annual spend            £412           £186

apparent effect: +£226 per customer  (+121%)

Members are obviously different people — they joined because they shop often. Model the propensity to join from pre-enrolment characteristics, then match:

BEFORE MATCHING              members    non-members    difference
prior-year orders               6.2           1.4         +4.8
prior-year spend             £358          £102          +£256
tenure (months)                28             11          +17
mobile app installed          71%            22%         +49pp

AFTER MATCHING (24,000 pairs)
prior-year orders               6.2           6.0         +0.2
prior-year spend             £358          £351           +£7
tenure (months)                28             27           +1
mobile app installed          71%            69%          +2pp
                                                    ↑ balanced

matched outcome comparison
avg annual spend             £412          £377          +£35

The estimated effect fell from £226 to £35. Most of the apparent value of membership was the customers who joined, not the programme. £35 is a defensible estimate of what to compare the programme’s cost against — Loyalty Programmes.

Step 3 is what makes this credible. The balance table is the output people should be shown, not the effect estimate alone. Unbalanced covariates after matching mean the matching failed — Confounding Variables.

The limitation that can’t be engineered away

Matching balances what you measured. Nothing else.

matched on:  prior orders, spend, tenure, app install, channel, region

NOT matched on:  how much the customer likes your brand
                 whether they were about to increase spending anyway
                 income change
                 whether a competitor closed nearby
                 the intent that made them join in the first place

and the LAST one is the whole problem — the decision to join
is caused by something you can't observe, and that something
also causes future spend

In plain terms: two customers with identical purchase histories, one of whom chose to join and one of whom didn’t, differ in exactly the unobserved way that drove the choice. Matching on history doesn’t make them the same person.

This is the fundamental difference from randomisation, and it’s a difference in kind rather than degree — Randomised Controlled Trials, Why Randomisation Works.

Doing it properly

  • Only pre-treatment variables in the model. Anything measured after enrolment is post-treatment and will bias the result — Multiple Regression
  • Check common support. If some treated units have propensity scores no untreated unit reaches, there’s nobody to compare them to. Either drop them — and say so, because you’ve changed the population your estimate applies to — or accept extrapolation
  • Always report the balance table. Standardised mean differences under about 0.1 is the usual working standard
  • Consider weighting instead of matching. Inverse probability weighting uses the whole sample rather than discarding unmatched units, and is generally more efficient
  • Run a sensitivity analysis. How strong would an unmeasured confounder have to be to overturn the result? If a modest one would do it, say that — it’s the most useful number in the whole analysis
  • Report it as an association with adjustment, not as a causal effect. The language matters and gets stripped as findings travel

When it’s the right tool

  • A programme already rolled out where no holdout was kept, and something must be said
  • Sizing an opportunity before committing to a proper test — matching gives a plausible range worth testing against
  • Where randomising is genuinely impossible — you can’t randomise who chooses to join a loyalty programme
  • As a sanity check on a test result, applied to observational data from the same period

The right response to a matched estimate is usually “so let’s test it.” £35 is a hypothesis with a number attached; a holdout on the next cohort would settle it — Holdout Groups, Incrementality Testing.

Where it interacts