Tags: experimentation concept

A-B Tests

Date: 2026-08-17


An A/B test is a controlled experiment run on live users of your own product: traffic is split at random between the current experience and one or more variants, and the difference in a pre-chosen metric is measured. The conventions around it — the vocabulary, the shape, the defaults — are what this note carries; the causal logic belongs to Controlled Experiments.


The shape of one

                      all eligible traffic
                              │
                    deterministic hash of user ID
                              │
              ┌───────────────┴───────────────┐
              │                               │
        control (A)                      variant (B)
        current experience               one change
              │                               │
        12,400 users                    12,390 users     ← check these match
        372 conversions                 421 conversions     Sample Ratio Mismatch
        3.00%                           3.40%
              └───────────────┬───────────────┘
                              │
                    +0.40pp absolute, +13.3% relative
                    95% CI: +1.2% to +25.9% relative

The ← check these match is the first thing to read, before any of the numbers below it: arm sizes that differ by more than chance mean the split didn’t work and nothing downstream is trustworthy — Sample Ratio Mismatch.

Absolute versus relative is the vocabulary error that causes the most confusion, including in meetings with people who should know better:

  • Absolute lift — 3.40% − 3.00% = +0.40 percentage points (pp)
  • Relative lift — 0.40 ÷ 3.00 = +13.3%

Both describe the same result. “Conversion went up 13%” and “conversion went up 0.4%” are both true and differ by a factor of 33. Say which one you mean, every time. The convention worth adopting: state effects in relative terms (it’s how MDE is usually specified) and state the underlying rates in absolute terms.

The vocabulary

TermMeans
Control / AThe current experience. Always include one, even when replacing something obviously broken
Variant / treatment / BThe changed experience. B, C, D… for more than one
ArmAny single group, control included
ExposureThe moment a user is assigned and could have seen the difference. Not page load
LiftThe difference in the primary metric, relative unless stated
Flat / inconclusiveThe interval includes zero — Inconclusive Results
Ship / roll backThe decision, which is not the same as “won / lost”

Defaults worth holding

  • One change per test, or you learn that something in a bundle worked. Bundling is legitimate when the bundle is what you’d ship — just know you can’t attribute within it
  • 50/50 split unless there’s a reason. Even splits maximise power for a given total sample — Traffic Allocation
  • User-level randomisation, sticky across sessions and devices where you can identify them — Randomisation Unit
  • One primary metric, chosen before launch — Overall Evaluation Criterion
  • Full business cycles, minimum one week, and never stop on a Wednesday because it crossed the line — Test Duration, Stopping Rules
  • Assignment logged as an event, or none of this is analysable — Experiment Assignment Tracking

What people expect that isn’t true

Most tests don’t win. Typical published win rates across mature programmes sit in the region of one in five to one in three, and lower for a well-optimised page. A programme reporting 70% winners has a measurement problem, not a talent advantage — Win Rate and Expected Value.

Effects are small. A genuine +2% relative on checkout conversion is a good result. The 30% lifts in vendor case studies are selected from thousands of tests and inflated by Winner’s Curse.

Underpowered tests are worse than no test. A test with 20% power to detect the effect you care about will usually return “inconclusive” when the effect is real, and when it does hit significance it will overstate the effect by a large factor. You’ve paid for traffic and bought misinformation — Statistical Power.

Significance is not a business case. A statistically clear +0.3% on a metric worth £40,000 a year is £120 and a permanent maintenance obligation — Practical vs Statistical Significance.

Where it’s the wrong tool

  • Not enough traffic. If the sample size for a plausible effect exceeds a quarter’s traffic, testing that page is theatre. Do research instead — Qualitative vs Quantitative Research
  • Users interfere with each other — marketplaces, two-sided platforms, anything with shared inventory. Assignment isn’t independent, so the arithmetic doesn’t hold — Switchback Tests, Geo Holdout Tests
  • The effect is long-run — brand, loyalty, churn over a year. A two-week test cannot see it — Holdout Groups
  • Nobody will act on either outcome. If it ships regardless, don’t spend the traffic

Where it interacts