Tags: experimentation concept
Controlled Experiments
Date: 2026-08-16
Two groups, differing in one thing, chosen at random. That’s the entire apparatus, and it’s the only method that tells you what a change caused rather than what it merely accompanied.
What it is
A controlled experiment compares two or more groups that are alike in every respect except the change under test, with membership decided at random.
The comparison group — the control — receives the existing experience. The treatment or variant receives the change. Everything else is held constant by the randomisation rather than by anyone’s effort.
An A/B test is one of these, not a synonym
The method is a century old — Fisher formalised randomised design for agricultural trials in the 1930s, medicine turned it into the randomised controlled trial, and direct mail was running split tests before the web existed. B testing is this method applied to a digital interface, with its own vocabulary and tooling. It is not a newer or better version.
The distinction is worth holding because the family is larger than the familiar case:
| Method | What gets randomised | Where |
|---|---|---|
| A-B Tests | Users | On your site |
| Geo Holdout Tests | Regions | Ad spend by area |
| Switchback Tests | Time blocks | Shared-resource systems |
| Holdout Groups | Users, permanently | Across the whole programme |
| Incrementality Testing | Users or geos | Inside ad platforms |
| Rollouts as Experiments | Users, during release | A ramp with a holdback |
All six are controlled experiments. Only the first is an A/B test.
Why this matters practically: at 12,000 visitors a week, an on-site A/B test is unaffordable for most changes — see the traffic constraint below. Knowing the general form is what lets you reach for a geo test or a year-long holdout and still make a causal claim, rather than concluding the question can’t be answered. It’s also the difference between “we can’t test this” and “we can’t A/B test this”.
Why it matters in a room: data scientists say “controlled experiment” and “treatment effect”; marketers say “A/B test” and “winner”. Same apparatus, different vocabulary, and recognising both is worth something.
Why it works when nothing else does
Every non-experimental comparison has the same flaw: the groups differ in ways you didn’t measure.
BEFORE / AFTER different weeks, campaigns, weather, mix
USERS WHO USED THE FEATURE self-selected — they were different already
MATCHED COHORTS matched on what you thought of, not what mattered
CONTROLLED EXPERIMENT differ only by the change, on average
Randomisation balances the variables you measured and the ones you never thought of — intent, mood, income, device age, whether they’d already decided to buy. No other method does the second part, and the second part is where the confounders live. See Why Randomisation Works and Confounding Variables.
In plain terms: you can’t list everything that might explain a difference between two groups. Randomising means you don’t have to — the unlisted things get split evenly too.
The four requirements
Break any one and the causal claim goes with it.
- Random assignment — not alternating, not by region, not “everyone except staff”. See Assignment and Bucketing
- Concurrent groups — both running at the same time. A sequential comparison is a before/after wearing a costume, and is confounded with everything that changed in between
- One difference — if the variant changes the layout and the copy and the price, a result tells you the bundle worked, not which part
- Outcomes measured identically in both arms — asymmetric tracking is the most common silent invalidator. See Experiment Assignment Tracking
What it can’t tell you
Worth being clear-eyed about, because the method’s authority gets overextended.
- Why. A test measures effect, never mechanism. The explanation is a hypothesis you brought with you, and if you didn’t bring one, a win teaches you nothing transferable — see Hypothesis Design
- Long-run effects. A two-week test measures two weeks. Habit formation, brand perception and repeat behaviour sit outside the window — see Holdout Groups
- Anything below its MDE. Most changes are smaller than most sites can detect
- Effects on people not in it. Spillover between arms — shared inventory, shared delivery capacity, users talking to each other — breaks the independence the maths assumes
- Whether it’s worth shipping. That’s cost, risk and margin, none of which are in the test
The ecommerce reality
The method is a century old and the constraint is arithmetic: it needs enough traffic. On 12,000 visitors a week at 3% conversion, detecting a 10% relative lift takes nine weeks — see Sample Size Calculation. That single fact shapes what a testing programme can be.
The honest consequence: most changes on most sites cannot be tested individually. That isn’t an argument against experimentation, it’s an argument for testing the changes big enough to measure, using held-back rollouts and holdouts for the rest, and being explicit about which decisions are being made on evidence and which on judgement.