Tags: statistics concept

Natural Experiments

Date: 2026-08-17


Situations where something outside your control assigned people to conditions in a way that’s as good as random. You get the causal leverage of an experiment without running one — provided you can defend the claim that the assignment had nothing to do with the outcome.


A natural experiment is a situation where something outside your control split people into groups in a way unrelated to the outcome, so comparing the groups supports a causal claim much as a randomised test would.

What makes one

The assignment must be arbitrary with respect to the thing you’re measuring. That’s the whole test, and it’s where most claimed natural experiments fail.

✓  a CDN outage affected one region for six hours
     → which region had the outage is unrelated to how likely
       its customers were to buy

✓  a pricing bug applied a wrong discount to orders placed
   in a 40-minute window
     → who ordered in that window is arbitrary

✓  free delivery threshold at £50 — customers at £49 vs £51
     → nearly identical people, discontinuously different treatment

✗  customers who received the email vs those who didn't
     → the send list was chosen. not arbitrary — Selection Bias

✗  users on the new app version vs the old
     → people who update quickly differ systematically

The two failing cases fail for the same reason: who ended up in each group was determined by something related to the outcome — Selection Bias.

The three shapes

Regression discontinuity. A threshold assigns treatment, and units just either side are comparable.

free delivery at £50

basket £47–£49.99    →  pays delivery      n = 8,400
basket £50–£52.99    →  free delivery      n = 11,200

completion rate      81.2%  vs  87.4%      +6.2pp

these customers are near-identical in intent and basket
composition. the only systematic difference is which side
of an arbitrary line they landed on

The caveat specific to this design: check for manipulation of the running variable. Customers can see the threshold and add items to cross it — which means the £50–£52.99 group contains people who deliberately topped up, and they’re different. A histogram of basket values showing a spike just above £50 is the diagnostic, and finding one usually invalidates the comparison — Shipping Thresholds.

Instrumental variables. Something affects treatment but has no other route to the outcome.

question   does faster delivery increase repeat purchase?
problem    customers who pay for express delivery are different

instrument  distance from the fulfilment centre
            → affects delivery speed
            → plausibly doesn't affect repeat purchase
              by any other route

use only the variation in delivery speed EXPLAINED by
distance, which is unrelated to customer type

The assumption is untestable and strong — that the instrument reaches the outcome only through the treatment. Here it’s arguable: distance might correlate with urban/rural, which affects shopping behaviour independently. That’s an exclusion-restriction violation and it’s the standard critique of any instrument.

Interrupted time series with a genuine shock. An outage, a regulatory change, a competitor’s failure — analysed as before/after with the shock as the intervention, usually via Difference-in-Differences against an unaffected group.

Worked: an outage as an experiment

a payment provider failed for 4 hours in one market

                        that market    comparable market
completion, outage hrs      41.2%           86.9%
completion, prior week
same hours                  85.8%           86.4%

DiD  =  (41.2 − 85.8) − (86.9 − 86.4)  =  −44.6 − 0.5  =  −45.1pp

In plain terms: the comparison market shows what “normal” drift looked like that week — half a point. Take that out, and the outage cost roughly 45 of every 100 checkouts that would otherwise have completed.

That’s a measurement you’d never be permitted to run deliberately — and it quantifies the revenue at risk from that provider, which is a real input to a redundancy decision. Incidents are natural experiments, and almost nobody analyses them as such.

The discipline that keeps it honest

The danger is that natural experiments are found after seeing an interesting pattern, which is The Garden of Forking Paths with better framing.

  • State the assignment mechanism first, in one sentence, and say why it’s arbitrary
  • Check pre-period comparability — the groups should look alike before the event
  • Run a placebo test on a period where nothing happened
  • Beware the shock changing several things at once. An outage affects speed, trust, and which customers were present. You’re measuring the bundle
  • Windows must be tight. The further from the event, the more else has changed
  • Report it as one estimate, not as proof. Natural experiments are stronger than correlation and weaker than an RCT, and saying so protects the finding when it’s challenged

Where they come from in practice

Worth keeping a list, because they’re perishable and only visible if someone’s looking:

outages and incidents            Incident Response postmortems
pricing and tagging bugs          the ones caught after some traffic
staggered rollouts                if the order was arbitrary — Progressive Delivery
platform-imposed changes          a browser update, a policy change
weather and external events       affecting some regions
supplier or competitor failures   stock-outs, closures
thresholds already in the product free delivery, loyalty tiers, discounts

Two of those rows are things you already produce: postmortems from Incident Response, and the ramp order in a Progressive Delivery rollout.

Every incident postmortem is a candidate. The data already exists, the counterfactual is unusually clean, and the analysis costs an afternoon — Annotation and Change Logs.

Where it interacts