Tags: statistics analytics concept
Correlation and Causation
Date: 2026-08-17
Everyone can recite the slogan and almost everyone acts on correlations anyway. The useful version isn’t the warning — it’s the specific list of five things that produce a correlation without causation, so you can name which one is the alternative explanation for the finding in front of you.
Correlation is two measures moving together; causation is one of them making the other move. The first is visible in data; the second is a claim about what would happen if you intervened.
The five rival explanations
For any observed correlation between A and B, at least six things could be true:
1 A → B A causes B ← the one you assumed
2 B → A reverse causation
3 C → A, C → B a confounder causes both
4 selection the correlation exists only in the data you collected
5 mediation A → M → B; the mechanism isn't what you think
6 coincidence with enough comparisons, some correlate by chance
The job is to say which of 2–6 you’ve ruled out and how. “Correlation isn’t causation” as a bare objection is unhelpful; naming the specific rival explanation is what moves a conversation forward.
Each, with the commerce version
Reverse causation. The arrow points the other way.
observed customers who use the wishlist have 2.4× the lifetime value
assumed wishlist → engagement → higher LTV
actual high-intent customers → use every feature, including wishlist
promoting the wishlist to everyone does not transfer the LTV
Confounding. A third variable drives both — the default explanation and the most common — Confounding Variables.
observed sessions using site search convert at 8.2%
sessions not using search convert at 2.1%
conclusion drawn "make search more prominent"
confounder purchase intent. people who know what they want
search for it AND buy it. search didn't cause
the intent, it accompanied it
Selection. The relationship is an artefact of who’s in the data.
observed app users churn less than web users
selection the app is installed by people already committed enough
to install an app. you're comparing committed customers
to everyone — Selection Bias, Survivorship Bias
Two related traps live here: the sample surviving to be measured at all — Selection Bias — and the units that dropped out being invisible — Survivorship Bias.
Mediation. A really does cause B, but through a path you’ve misidentified — which matters because it changes what to build.
observed adding delivery estimates raised conversion
assumed information → confidence → purchase
actual the estimate module pushed the price below the fold,
reducing price salience
building more information modules won't replicate it;
the effect was about layout
Coincidence. Test enough pairs and some correlate. Ten metrics give 10 × 9 ÷ 2 = 45 pairs; at 95% confidence each pair has a 5% false-positive rate, so 45 × 0.05 ≈ 2 look significant from nothing.
In plain terms: compare enough things and a couple will line up by pure luck, and they’ll look exactly like real relationships — The Multiple Comparisons Problem, The Garden of Forking Paths.
The tell that catches most of them
Ask what the counterfactual population looks like.
claim "email subscribers spend 3× more, so grow the list"
question: who are the people who WOULD subscribe if we pushed harder?
are they like current subscribers, or like the people who
have already declined to subscribe?
answer: the second. and the reason current subscribers spend more
is largely that they were already your best customers when
they subscribed.
the effect of ADDING a marginal subscriber is not the difference
between existing subscribers and non-subscribers
— Counterfactuals
That question is the whole of Counterfactuals, compressed.
This single question — “what would the people I’d affect have done otherwise?” — dissolves most correlational claims in commerce. The comparison group you have is almost never the comparison group your intervention would create.
What establishes causation
In descending order of how much you can trust it:
| Method | Handles unknown confounders? | Cost |
|---|---|---|
| Randomised Controlled Trials / A/B test | Yes — that’s the point | Traffic, time |
| Natural Experiments | Mostly, if the assignment really was as-good-as-random | Requires the world to cooperate |
| Difference-in-Differences | Only those constant over time | Needs a valid comparison group |
| Propensity Score Matching | No — observed confounders only | Cheap, and routinely oversold |
| Multiple Regression controls | No — observed only, and can make things worse | Cheapest, weakest |
| Correlation alone | No | Free |
The bottom three all share one limitation: they adjust for confounders you measured and thought of. The value of randomisation is that it balances the ones you didn’t — including the ones nobody has ever named — Why Randomisation Works.
When to act on a correlation anyway
Being absolutist about this is its own failure. Acting without proof is fine when:
- The action is cheap and reversible. Testing costs traffic; sometimes just doing it is cheaper than measuring it
- A mechanism is plausible and specific. “Faster pages convert better” has physical reasoning behind it and replicates widely — Performance and Conversion
- It’s a prediction, not an intervention. A model predicting churn doesn’t need causal validity to be useful for targeting — it needs to predict. Causation matters when you intend to change something
- You’ll test it next. Correlational findings are excellent hypothesis generators, and that’s their proper role — Hypothesis Design, Path Analysis
The failure isn’t using correlations. It’s spending money on one while describing it as causal, and then attributing the result to the wrong mechanism.
Where it interacts
- Correlation — measuring the association this note is about interpreting
- Confounding Variables — the single most common of the five explanations
- Counterfactuals — the comparison a causal claim is really making, and the question that exposes bad ones
- Regression to the Mean — a sixth pattern that looks like causation, where the intervention gets credit for a statistical artefact