Tags: statistics concept

Randomised Controlled Trials

Date: 2026-08-17


The design where units are assigned to treatment or control by chance. Its unique property is that it balances confounders you never thought of — which is why it’s the only method whose validity doesn’t depend on having correctly guessed what else matters.


A randomised controlled trial (RCT) assigns units to treatment or control by chance, then compares outcomes between the groups. An A/B test is one.

What randomisation buys

Every other causal method adjusts for a list of variables someone wrote down. Randomisation doesn’t need the list.

OBSERVATIONAL ADJUSTMENT             RANDOMISATION

"we controlled for device,           assignment is independent of
 channel, and new/returning"          EVERY pre-existing characteristic
                                      by construction
what about intent? urgency?
income? whether they'd already       so the groups are comparable on
decided? mood? the fact that         intent, urgency, income, mood,
they'd just read a review?            and every variable nobody has
                                      ever named
each unmeasured confounder
remains

In plain terms: you don’t have to know what confounds a comparison in order to eliminate it. That is a genuinely remarkable property and it is the reason experimentation dominates every alternative where it’s affordable — Confounding Variables, Why Randomisation Works.

Note it’s a statement about averages, not about any single trial. One randomisation can land unbalanced by chance; the procedure is unbiased. This is why checking baseline balance is worth doing, and why a large imbalance is a signal something’s broken rather than unlucky — Sample Ratio Mismatch.

The four requirements

1  RANDOM ASSIGNMENT       by a process independent of any characteristic
                           of the unit. deterministic hashing of an ID
                           counts; "alternate visitors" does not

2  A CONCURRENT CONTROL    running at the same time, not last month.
                           a before/after comparison is not an RCT

3  ADEQUATE SIZE           enough units for the difference to be
                           distinguishable from noise
                           — Statistical Power

4  ANALYSIS AS ASSIGNED    every unit counted in the arm it was
                           assigned to, whatever happened next
                           ← see below

Requirement 3 is where most commercial tests fail before they start, and it’s decided at design time rather than discovered at analysis — Statistical Power.

Requirement 4 is the one broken most often.

Intention to treat

Analyse people by the group they were assigned to, not by what they actually experienced.

variant assigned      10,000 users
  of whom saw it       8,200
  didn't see it        1,800   (left before the element rendered,
                                 blocked the script, bounced instantly)

✗ ANALYSING THE 8,200
    those 1,800 are not random — they're the fastest bouncers,
    the least engaged, the ones on poor connections.
    removing them from the variant but not from control makes
    the arms non-comparable, and the randomisation is gone

✓ ANALYSING ALL 10,000
    preserves comparability. dilutes the measured effect,
    but the estimate is unbiased

The trade is real: intention to treat dilutes the effect towards zero in proportion to the non-exposed share. The fix isn’t to filter afterwards — it’s to assign at the point of exposure, so the non-exposed were never randomised in the first place. Filtering after assignment is Selection Bias; assigning later is design — Sample Pollution.

The one safe exclusion is on characteristics fixed independently of treatment — bot user-agents, internal IPs, assignment before QA sign-off.

What an RCT still can’t tell you

Being clear about the limits, because “we randomised” gets treated as a universal warrant:

  • Whether it generalises. The effect is estimated for the population, period and context you tested in. A Black Friday result may not hold in March — Seasonality in Tests
  • Long-run effects. Two weeks cannot see a change in annual retention — Holdout Groups
  • Why. Randomisation establishes that the change caused the difference, not through what mechanism — Secondary and Diagnostic Metrics
  • Anything, if units interfere. SUTVA — the stable unit treatment value assumption, that one unit’s treatment doesn’t affect another’s outcome — fails with shared inventory, marketplaces and social features. Randomisation is intact; the interpretation isn’t — Switchback Tests
  • Effects it wasn’t powered for. A null from an underpowered trial is not evidence of no effect — Inconclusive Results

The vocabulary, since it travels

The term comes from clinical research, and some of its apparatus transfers usefully:

  • Blinding — subjects and assessors not knowing the arm. Users are naturally blind in web tests; analysts frequently aren’t, and an analyst who knows which arm is “the new one” makes systematically different judgement calls. Blinded analysis is cheap and rarely done
  • Pre-registration — the analysis plan filed before data collection. Standard in trials, still uncommon in commercial testing, and the thing that makes a p-value mean anything — Pre-Registration
  • Stratified randomisation — balancing on a known important variable rather than trusting chance, which reduces variance at no cost — Sampling Methods
  • Cluster randomisation — assigning whole groups where individuals interfere

Where it interacts

  • Controlled Experiments — the same idea in this vault’s experimentation vocabulary; that note carries the practice, this one the statistical grounding
  • Counterfactuals — what randomisation constructs, and why it’s the only method that does so without assumptions
  • Natural Experiments — randomisation the world performed, with the same benefit and a weaker guarantee
  • Ethics of Experimentation — the constraint on when randomising people is acceptable