Tags: statistics concept
Propensity Score Matching
Date: 2026-08-17
Building a comparison group from observational data by pairing treated units with untreated ones that look similar. It’s genuinely useful and routinely oversold — it can only balance the variables you measured, so it fixes the confounding you already knew about and does nothing for the confounding that was the actual problem.
Propensity score matching estimates each unit’s probability of having been treated from its characteristics — its propensity score — and compares treated units with untreated ones of similar score.
The method
1 MODEL estimate P(treated | characteristics) for every unit
— usually Logistic Regression
this probability is the PROPENSITY SCORE
2 MATCH pair each treated unit with untreated unit(s) of
similar score — nearest neighbour, calipers, or
weight the whole sample by it
3 CHECK verify the matched groups now look alike on every
covariate ← the step that gets skipped
4 COMPARE difference in outcome between matched groups
The insight that makes it work: you don’t need to match on twenty variables individually. Units with the same probability of being treated have, on average, the same distribution of all the variables that went into computing it. One number stands in for the whole set.
Worked
Loyalty programme members appear to be worth far more:
members non-members
24,000 180,000
avg annual spend £412 £186
apparent effect: +£226 per customer (+121%)
Members are obviously different people — they joined because they shop often. Model the propensity to join from pre-enrolment characteristics, then match:
BEFORE MATCHING members non-members difference
prior-year orders 6.2 1.4 +4.8
prior-year spend £358 £102 +£256
tenure (months) 28 11 +17
mobile app installed 71% 22% +49pp
AFTER MATCHING (24,000 pairs)
prior-year orders 6.2 6.0 +0.2
prior-year spend £358 £351 +£7
tenure (months) 28 27 +1
mobile app installed 71% 69% +2pp
↑ balanced
matched outcome comparison
avg annual spend £412 £377 +£35
The estimated effect fell from £226 to £35. Most of the apparent value of membership was the customers who joined, not the programme. £35 is a defensible estimate of what to compare the programme’s cost against — Loyalty Programmes.
Step 3 is what makes this credible. The balance table is the output people should be shown, not the effect estimate alone. Unbalanced covariates after matching mean the matching failed — Confounding Variables.
The limitation that can’t be engineered away
Matching balances what you measured. Nothing else.
matched on: prior orders, spend, tenure, app install, channel, region
NOT matched on: how much the customer likes your brand
whether they were about to increase spending anyway
income change
whether a competitor closed nearby
the intent that made them join in the first place
and the LAST one is the whole problem — the decision to join
is caused by something you can't observe, and that something
also causes future spend
In plain terms: two customers with identical purchase histories, one of whom chose to join and one of whom didn’t, differ in exactly the unobserved way that drove the choice. Matching on history doesn’t make them the same person.
This is the fundamental difference from randomisation, and it’s a difference in kind rather than degree — Randomised Controlled Trials, Why Randomisation Works.
Doing it properly
- Only pre-treatment variables in the model. Anything measured after enrolment is post-treatment and will bias the result — Multiple Regression
- Check common support. If some treated units have propensity scores no untreated unit reaches, there’s nobody to compare them to. Either drop them — and say so, because you’ve changed the population your estimate applies to — or accept extrapolation
- Always report the balance table. Standardised mean differences under about 0.1 is the usual working standard
- Consider weighting instead of matching. Inverse probability weighting uses the whole sample rather than discarding unmatched units, and is generally more efficient
- Run a sensitivity analysis. How strong would an unmeasured confounder have to be to overturn the result? If a modest one would do it, say that — it’s the most useful number in the whole analysis
- Report it as an association with adjustment, not as a causal effect. The language matters and gets stripped as findings travel
When it’s the right tool
- A programme already rolled out where no holdout was kept, and something must be said
- Sizing an opportunity before committing to a proper test — matching gives a plausible range worth testing against
- Where randomising is genuinely impossible — you can’t randomise who chooses to join a loyalty programme
- As a sanity check on a test result, applied to observational data from the same period
The right response to a matched estimate is usually “so let’s test it.” £35 is a hypothesis with a number attached; a holdout on the next cohort would settle it — Holdout Groups, Incrementality Testing.
Where it interacts
- Counterfactuals — matching constructs one from lookalikes, and inherits the limits of that
- Confounding Variables — the observed ones this addresses, and the unobserved ones it can’t
- Logistic Regression — the model that produces the score
- Selection Bias — the thing being adjusted for, and the reason the adjustment is incomplete