Tags: statistics concept
Natural Experiments
Date: 2026-08-17
Situations where something outside your control assigned people to conditions in a way that’s as good as random. You get the causal leverage of an experiment without running one — provided you can defend the claim that the assignment had nothing to do with the outcome.
A natural experiment is a situation where something outside your control split people into groups in a way unrelated to the outcome, so comparing the groups supports a causal claim much as a randomised test would.
What makes one
The assignment must be arbitrary with respect to the thing you’re measuring. That’s the whole test, and it’s where most claimed natural experiments fail.
✓ a CDN outage affected one region for six hours
→ which region had the outage is unrelated to how likely
its customers were to buy
✓ a pricing bug applied a wrong discount to orders placed
in a 40-minute window
→ who ordered in that window is arbitrary
✓ free delivery threshold at £50 — customers at £49 vs £51
→ nearly identical people, discontinuously different treatment
✗ customers who received the email vs those who didn't
→ the send list was chosen. not arbitrary — Selection Bias
✗ users on the new app version vs the old
→ people who update quickly differ systematically
The two failing cases fail for the same reason: who ended up in each group was determined by something related to the outcome — Selection Bias.
The three shapes
Regression discontinuity. A threshold assigns treatment, and units just either side are comparable.
free delivery at £50
basket £47–£49.99 → pays delivery n = 8,400
basket £50–£52.99 → free delivery n = 11,200
completion rate 81.2% vs 87.4% +6.2pp
these customers are near-identical in intent and basket
composition. the only systematic difference is which side
of an arbitrary line they landed on
The caveat specific to this design: check for manipulation of the running variable. Customers can see the threshold and add items to cross it — which means the £50–£52.99 group contains people who deliberately topped up, and they’re different. A histogram of basket values showing a spike just above £50 is the diagnostic, and finding one usually invalidates the comparison — Shipping Thresholds.
Instrumental variables. Something affects treatment but has no other route to the outcome.
question does faster delivery increase repeat purchase?
problem customers who pay for express delivery are different
instrument distance from the fulfilment centre
→ affects delivery speed
→ plausibly doesn't affect repeat purchase
by any other route
use only the variation in delivery speed EXPLAINED by
distance, which is unrelated to customer type
The assumption is untestable and strong — that the instrument reaches the outcome only through the treatment. Here it’s arguable: distance might correlate with urban/rural, which affects shopping behaviour independently. That’s an exclusion-restriction violation and it’s the standard critique of any instrument.
Interrupted time series with a genuine shock. An outage, a regulatory change, a competitor’s failure — analysed as before/after with the shock as the intervention, usually via Difference-in-Differences against an unaffected group.
Worked: an outage as an experiment
a payment provider failed for 4 hours in one market
that market comparable market
completion, outage hrs 41.2% 86.9%
completion, prior week
same hours 85.8% 86.4%
DiD = (41.2 − 85.8) − (86.9 − 86.4) = −44.6 − 0.5 = −45.1pp
In plain terms: the comparison market shows what “normal” drift looked like that week — half a point. Take that out, and the outage cost roughly 45 of every 100 checkouts that would otherwise have completed.
That’s a measurement you’d never be permitted to run deliberately — and it quantifies the revenue at risk from that provider, which is a real input to a redundancy decision. Incidents are natural experiments, and almost nobody analyses them as such.
The discipline that keeps it honest
The danger is that natural experiments are found after seeing an interesting pattern, which is The Garden of Forking Paths with better framing.
- State the assignment mechanism first, in one sentence, and say why it’s arbitrary
- Check pre-period comparability — the groups should look alike before the event
- Run a placebo test on a period where nothing happened
- Beware the shock changing several things at once. An outage affects speed, trust, and which customers were present. You’re measuring the bundle
- Windows must be tight. The further from the event, the more else has changed
- Report it as one estimate, not as proof. Natural experiments are stronger than correlation and weaker than an RCT, and saying so protects the finding when it’s challenged
Where they come from in practice
Worth keeping a list, because they’re perishable and only visible if someone’s looking:
outages and incidents Incident Response postmortems
pricing and tagging bugs the ones caught after some traffic
staggered rollouts if the order was arbitrary — Progressive Delivery
platform-imposed changes a browser update, a policy change
weather and external events affecting some regions
supplier or competitor failures stock-outs, closures
thresholds already in the product free delivery, loyalty tiers, discounts
Two of those rows are things you already produce: postmortems from Incident Response, and the ramp order in a Progressive Delivery rollout.
Every incident postmortem is a candidate. The data already exists, the counterfactual is unusually clean, and the analysis costs an afternoon — Annotation and Change Logs.
Where it interacts
- Counterfactuals — natural experiments supply an unusually credible one
- Difference-in-Differences — the analysis method most often applied to them
- Randomised Controlled Trials — the deliberate version, with a guarantee rather than an argument
- Correlation and Causation — natural experiments are the main route to a causal claim when testing isn’t possible