Tags: experimentation concept
A-B Tests
Date: 2026-08-17
An A/B test is a controlled experiment run on live users of your own product: traffic is split at random between the current experience and one or more variants, and the difference in a pre-chosen metric is measured. The conventions around it — the vocabulary, the shape, the defaults — are what this note carries; the causal logic belongs to Controlled Experiments.
The shape of one
all eligible traffic
│
deterministic hash of user ID
│
┌───────────────┴───────────────┐
│ │
control (A) variant (B)
current experience one change
│ │
12,400 users 12,390 users ← check these match
372 conversions 421 conversions Sample Ratio Mismatch
3.00% 3.40%
└───────────────┬───────────────┘
│
+0.40pp absolute, +13.3% relative
95% CI: +1.2% to +25.9% relative
The ← check these match is the first thing to read, before any of the numbers below it: arm sizes that differ by more than chance mean the split didn’t work and nothing downstream is trustworthy — Sample Ratio Mismatch.
Absolute versus relative is the vocabulary error that causes the most confusion, including in meetings with people who should know better:
- Absolute lift — 3.40% − 3.00% = +0.40 percentage points (pp)
- Relative lift — 0.40 ÷ 3.00 = +13.3%
Both describe the same result. “Conversion went up 13%” and “conversion went up 0.4%” are both true and differ by a factor of 33. Say which one you mean, every time. The convention worth adopting: state effects in relative terms (it’s how MDE is usually specified) and state the underlying rates in absolute terms.
The vocabulary
| Term | Means |
|---|---|
| Control / A | The current experience. Always include one, even when replacing something obviously broken |
| Variant / treatment / B | The changed experience. B, C, D… for more than one |
| Arm | Any single group, control included |
| Exposure | The moment a user is assigned and could have seen the difference. Not page load |
| Lift | The difference in the primary metric, relative unless stated |
| Flat / inconclusive | The interval includes zero — Inconclusive Results |
| Ship / roll back | The decision, which is not the same as “won / lost” |
Defaults worth holding
- One change per test, or you learn that something in a bundle worked. Bundling is legitimate when the bundle is what you’d ship — just know you can’t attribute within it
- 50/50 split unless there’s a reason. Even splits maximise power for a given total sample — Traffic Allocation
- User-level randomisation, sticky across sessions and devices where you can identify them — Randomisation Unit
- One primary metric, chosen before launch — Overall Evaluation Criterion
- Full business cycles, minimum one week, and never stop on a Wednesday because it crossed the line — Test Duration, Stopping Rules
- Assignment logged as an event, or none of this is analysable — Experiment Assignment Tracking
What people expect that isn’t true
Most tests don’t win. Typical published win rates across mature programmes sit in the region of one in five to one in three, and lower for a well-optimised page. A programme reporting 70% winners has a measurement problem, not a talent advantage — Win Rate and Expected Value.
Effects are small. A genuine +2% relative on checkout conversion is a good result. The 30% lifts in vendor case studies are selected from thousands of tests and inflated by Winner’s Curse.
Underpowered tests are worse than no test. A test with 20% power to detect the effect you care about will usually return “inconclusive” when the effect is real, and when it does hit significance it will overstate the effect by a large factor. You’ve paid for traffic and bought misinformation — Statistical Power.
Significance is not a business case. A statistically clear +0.3% on a metric worth £40,000 a year is £120 and a permanent maintenance obligation — Practical vs Statistical Significance.
Where it’s the wrong tool
- Not enough traffic. If the sample size for a plausible effect exceeds a quarter’s traffic, testing that page is theatre. Do research instead — Qualitative vs Quantitative Research
- Users interfere with each other — marketplaces, two-sided platforms, anything with shared inventory. Assignment isn’t independent, so the arithmetic doesn’t hold — Switchback Tests, Geo Holdout Tests
- The effect is long-run — brand, loyalty, churn over a year. A two-week test cannot see it — Holdout Groups
- Nobody will act on either outcome. If it ships regardless, don’t spend the traffic
Where it interacts
- Controlled Experiments — the parent form, and why randomisation is what makes any of this causal
- Reading a Test Result — the ordered checks before believing a number
- Experiment QA — the checks before traffic, which catch more bad tests than any analysis does
- Client-Side vs Server-Side Testing — where the variant decision happens, and everything downstream of that choice