Tags: statistics concept

Counterfactuals

Date: 2026-08-17


The comparison every causal claim is secretly making: what would have happened to these same people, at the same time, if you hadn’t done it. It’s never observable — you only get one version of events per person — which is why causal inference is entirely a business of constructing substitutes for something you can’t see.


A counterfactual is the outcome that would have happened to the same units, at the same time, without the intervention — the unobserved half of every causal comparison.

The fundamental problem

for one customer, two potential outcomes exist

  Y(1)   what they'd do if they saw the variant
  Y(0)   what they'd do if they saw the control

the causal effect for them is  Y(1) − Y(0)

you observe EXACTLY ONE of these. ever.
the other is permanently missing.

This is the fundamental problem of causal inference, and it’s not a data limitation you can spend your way out of. No amount of instrumentation reveals what a person who saw the variant would have done had they not.

In plain terms: you can never measure a causal effect on an individual. You can only estimate an average effect across a group, by finding people to stand in for the missing halves.

What each method uses as the stand-in

Every causal method is a different answer to “what plays the role of the unobserved counterfactual?”

MethodThe counterfactual is…Holds when
Randomised Controlled TrialsThe randomly assigned control groupAlways — randomisation makes the groups exchangeable
Difference-in-DifferencesThe treated group’s own trend, borrowed from an untreated groupParallel trends
Propensity Score MatchingSimilar-looking untreated unitsAll confounders observed
Natural ExperimentsUnits on the other side of an as-good-as-random boundaryThe assignment really was arbitrary
Before/afterThe same units, earlierNothing else changed — almost never true
Holdout GroupsA permanently untreated sliceMaintained isolation over time

The before/after row is the one used most and justified least. “Conversion was 3.1% before the redesign and 3.4% after” compares two different time periods with different traffic, weather, campaigns and competitors — the counterfactual assumed is “everything else would have stayed identical”, which is never true.

Why randomisation solves it

Randomisation doesn’t reveal any individual’s counterfactual. It does something subtler and sufficient: it makes the control group’s average outcome an unbiased estimate of what the treatment group would have done untreated.

randomly split 200,000 users

group A and group B are, on average, identical in every
characteristic — measured, unmeasured, and unimagined —
because assignment was independent of all of them

so:  average Y(0) for group B  ≈  average Y(0) for group A
     and you OBSERVE the latter

     estimated effect  =  mean(B observed)  −  mean(A observed)

The groups are exchangeable — you could swap which one got treated and expect the same answer. That’s the property everything else on this page is trying to approximate — Why Randomisation Works.

The question that exposes bad claims

Asking “what’s the counterfactual?” out loud dissolves most weak causal reasoning.

claim  "customers who use the wishlist have 2.4× the lifetime value,
        so we should push wishlist adoption"

counterfactual needed
       what would the wishlist users have spent WITHOUT the wishlist?

the comparison being made
       what did NON-users spend?

these are different questions. non-users are different people —
lower intent, less engaged — so they're a poor stand-in for
"wishlist users, counterfactually deprived of the wishlist"
claim  "the campaign generated £400,000 in attributed revenue"

counterfactual needed
       how much of that would have happened anyway?

for branded search and retargeting, frequently most of it —
which is what Incrementality Testing measures and attribution
does not

“Compared to what?” is the entire discipline in three words.

Which effect you’re estimating

Worth distinguishing, because they differ and get conflated:

ATE   average treatment effect
      the effect if EVERYONE were treated

ATT   average treatment effect on the treated
      the effect on those who actually got it
      ← what matching and DiD usually estimate

LATE  local average treatment effect
      the effect on those whose behaviour the assignment
      actually changed
      ← what an instrument or a discontinuity gives you

They can differ a lot. An email’s effect on people who opted in (ATT) tells you little about its effect if sent to everyone (ATE), because the opt-in group is more receptive. Rolling a campaign out to the whole base on the strength of an ATT estimate is one of the commonest overreaches in growth work.

Practical use

  • Write the counterfactual down before any causal claim. One sentence: “what these people would have done otherwise, estimated by ___”
  • Name the assumption that makes your stand-in valid, and say what would break it
  • Treat before/after as a prompt, never as evidence — Annotation and Change Logs
  • Prefer a holdout to a reconstruction. Keeping 5% untreated costs 5% of the benefit and buys a real counterfactual, permanently — usually a better deal than modelling one afterwards

Where it interacts