Tags: statistics concept
Counterfactuals
Date: 2026-08-17
The comparison every causal claim is secretly making: what would have happened to these same people, at the same time, if you hadn’t done it. It’s never observable — you only get one version of events per person — which is why causal inference is entirely a business of constructing substitutes for something you can’t see.
A counterfactual is the outcome that would have happened to the same units, at the same time, without the intervention — the unobserved half of every causal comparison.
The fundamental problem
for one customer, two potential outcomes exist
Y(1) what they'd do if they saw the variant
Y(0) what they'd do if they saw the control
the causal effect for them is Y(1) − Y(0)
you observe EXACTLY ONE of these. ever.
the other is permanently missing.
This is the fundamental problem of causal inference, and it’s not a data limitation you can spend your way out of. No amount of instrumentation reveals what a person who saw the variant would have done had they not.
In plain terms: you can never measure a causal effect on an individual. You can only estimate an average effect across a group, by finding people to stand in for the missing halves.
What each method uses as the stand-in
Every causal method is a different answer to “what plays the role of the unobserved counterfactual?”
| Method | The counterfactual is… | Holds when |
|---|---|---|
| Randomised Controlled Trials | The randomly assigned control group | Always — randomisation makes the groups exchangeable |
| Difference-in-Differences | The treated group’s own trend, borrowed from an untreated group | Parallel trends |
| Propensity Score Matching | Similar-looking untreated units | All confounders observed |
| Natural Experiments | Units on the other side of an as-good-as-random boundary | The assignment really was arbitrary |
| Before/after | The same units, earlier | Nothing else changed — almost never true |
| Holdout Groups | A permanently untreated slice | Maintained isolation over time |
The before/after row is the one used most and justified least. “Conversion was 3.1% before the redesign and 3.4% after” compares two different time periods with different traffic, weather, campaigns and competitors — the counterfactual assumed is “everything else would have stayed identical”, which is never true.
Why randomisation solves it
Randomisation doesn’t reveal any individual’s counterfactual. It does something subtler and sufficient: it makes the control group’s average outcome an unbiased estimate of what the treatment group would have done untreated.
randomly split 200,000 users
group A and group B are, on average, identical in every
characteristic — measured, unmeasured, and unimagined —
because assignment was independent of all of them
so: average Y(0) for group B ≈ average Y(0) for group A
and you OBSERVE the latter
estimated effect = mean(B observed) − mean(A observed)
The groups are exchangeable — you could swap which one got treated and expect the same answer. That’s the property everything else on this page is trying to approximate — Why Randomisation Works.
The question that exposes bad claims
Asking “what’s the counterfactual?” out loud dissolves most weak causal reasoning.
claim "customers who use the wishlist have 2.4× the lifetime value,
so we should push wishlist adoption"
counterfactual needed
what would the wishlist users have spent WITHOUT the wishlist?
the comparison being made
what did NON-users spend?
these are different questions. non-users are different people —
lower intent, less engaged — so they're a poor stand-in for
"wishlist users, counterfactually deprived of the wishlist"
claim "the campaign generated £400,000 in attributed revenue"
counterfactual needed
how much of that would have happened anyway?
for branded search and retargeting, frequently most of it —
which is what Incrementality Testing measures and attribution
does not
“Compared to what?” is the entire discipline in three words.
Which effect you’re estimating
Worth distinguishing, because they differ and get conflated:
ATE average treatment effect
the effect if EVERYONE were treated
ATT average treatment effect on the treated
the effect on those who actually got it
← what matching and DiD usually estimate
LATE local average treatment effect
the effect on those whose behaviour the assignment
actually changed
← what an instrument or a discontinuity gives you
They can differ a lot. An email’s effect on people who opted in (ATT) tells you little about its effect if sent to everyone (ATE), because the opt-in group is more receptive. Rolling a campaign out to the whole base on the strength of an ATT estimate is one of the commonest overreaches in growth work.
Practical use
- Write the counterfactual down before any causal claim. One sentence: “what these people would have done otherwise, estimated by ___”
- Name the assumption that makes your stand-in valid, and say what would break it
- Treat before/after as a prompt, never as evidence — Annotation and Change Logs
- Prefer a holdout to a reconstruction. Keeping 5% untreated costs 5% of the benefit and buys a real counterfactual, permanently — usually a better deal than modelling one afterwards
Where it interacts
- Randomised Controlled Trials — the method that constructs a valid counterfactual by design
- Correlation and Causation — the counterfactual question is the practical test for the five failure modes listed there
- Difference-in-Differences and Natural Experiments — the two most useful approximations when randomising isn’t available
- Incrementality Testing — the commerce application, where the counterfactual is the entire commercial question