Tags: experimentation concept

Controlled Experiments

Date: 2026-08-16


Two groups, differing in one thing, chosen at random. That’s the entire apparatus, and it’s the only method that tells you what a change caused rather than what it merely accompanied.


What it is

A controlled experiment compares two or more groups that are alike in every respect except the change under test, with membership decided at random.

The comparison group — the control — receives the existing experience. The treatment or variant receives the change. Everything else is held constant by the randomisation rather than by anyone’s effort.

An A/B test is one of these, not a synonym

The method is a century old — Fisher formalised randomised design for agricultural trials in the 1930s, medicine turned it into the randomised controlled trial, and direct mail was running split tests before the web existed. B testing is this method applied to a digital interface, with its own vocabulary and tooling. It is not a newer or better version.

The distinction is worth holding because the family is larger than the familiar case:

MethodWhat gets randomisedWhere
A-B TestsUsersOn your site
Geo Holdout TestsRegionsAd spend by area
Switchback TestsTime blocksShared-resource systems
Holdout GroupsUsers, permanentlyAcross the whole programme
Incrementality TestingUsers or geosInside ad platforms
Rollouts as ExperimentsUsers, during releaseA ramp with a holdback

All six are controlled experiments. Only the first is an A/B test.

Why this matters practically: at 12,000 visitors a week, an on-site A/B test is unaffordable for most changes — see the traffic constraint below. Knowing the general form is what lets you reach for a geo test or a year-long holdout and still make a causal claim, rather than concluding the question can’t be answered. It’s also the difference between “we can’t test this” and “we can’t A/B test this”.

Why it matters in a room: data scientists say “controlled experiment” and “treatment effect”; marketers say “A/B test” and “winner”. Same apparatus, different vocabulary, and recognising both is worth something.

Why it works when nothing else does

Every non-experimental comparison has the same flaw: the groups differ in ways you didn’t measure.

BEFORE / AFTER              different weeks, campaigns, weather, mix
USERS WHO USED THE FEATURE  self-selected — they were different already
MATCHED COHORTS             matched on what you thought of, not what mattered
CONTROLLED EXPERIMENT       differ only by the change, on average

Randomisation balances the variables you measured and the ones you never thought of — intent, mood, income, device age, whether they’d already decided to buy. No other method does the second part, and the second part is where the confounders live. See Why Randomisation Works and Confounding Variables.

In plain terms: you can’t list everything that might explain a difference between two groups. Randomising means you don’t have to — the unlisted things get split evenly too.

The four requirements

Break any one and the causal claim goes with it.

  1. Random assignment — not alternating, not by region, not “everyone except staff”. See Assignment and Bucketing
  2. Concurrent groups — both running at the same time. A sequential comparison is a before/after wearing a costume, and is confounded with everything that changed in between
  3. One difference — if the variant changes the layout and the copy and the price, a result tells you the bundle worked, not which part
  4. Outcomes measured identically in both arms — asymmetric tracking is the most common silent invalidator. See Experiment Assignment Tracking

What it can’t tell you

Worth being clear-eyed about, because the method’s authority gets overextended.

  • Why. A test measures effect, never mechanism. The explanation is a hypothesis you brought with you, and if you didn’t bring one, a win teaches you nothing transferable — see Hypothesis Design
  • Long-run effects. A two-week test measures two weeks. Habit formation, brand perception and repeat behaviour sit outside the window — see Holdout Groups
  • Anything below its MDE. Most changes are smaller than most sites can detect
  • Effects on people not in it. Spillover between arms — shared inventory, shared delivery capacity, users talking to each other — breaks the independence the maths assumes
  • Whether it’s worth shipping. That’s cost, risk and margin, none of which are in the test

The ecommerce reality

The method is a century old and the constraint is arithmetic: it needs enough traffic. On 12,000 visitors a week at 3% conversion, detecting a 10% relative lift takes nine weeks — see Sample Size Calculation. That single fact shapes what a testing programme can be.

The honest consequence: most changes on most sites cannot be tested individually. That isn’t an argument against experimentation, it’s an argument for testing the changes big enough to measure, using held-back rollouts and holdouts for the rest, and being explicit about which decisions are being made on evidence and which on judgement.