Tags: experimentation commerce concept
Holdout Groups
Date: 2026-08-16
A slice of users kept permanently untreated, so you can measure what a year of shipped changes was actually worth in total. It exists because individually significant wins routinely sum to more than the business grew — and the holdout is the only thing that says so.
What it is
A holdout group is a randomly selected portion of traffic or customers excluded from treatment — from one programme, or from everything — and left that way for a long period.
Unlike a test control, which ends when the test does, a holdout persists. That’s the whole mechanism: it measures the cumulative effect of everything shipped, over a horizon no individual test can see.
The problem it solves
Sum the wins from a year of testing and compare to what the business did:
claimed from 14 winning tests +18% conversion
actual year-on-year change +3% conversion
That gap is normal, and it’s the single most common credibility problem in a testing programme. It has several causes at once:
- Winner’s Curse — measured effects of winners are biased upward, systematically
- Effects decay. Novelty and Primacy Effects fade; some wins are temporary by construction
- Effects overlap. Two tests that each move the same behaviour don’t add
- The counterfactual moved. The baseline was never going to stay still — seasonality, mix, competitors
- Some “wins” were false positives, which is what a 5% error rate means in practice — The Multiple Comparisons Problem
A holdout measures the sum directly and needs none of that adjusted for.
Two kinds
| Type | Excluded from | Answers |
|---|---|---|
| Programme holdout | Every experiment-driven change | What is the testing programme worth? |
| Channel holdout | A marketing channel or campaign | Is this spend incremental? — Incrementality Testing |
The channel version is a different tool answering an adjacent question, and it’s usually a geo split rather than a user split because ad platforms can’t reliably exclude individuals — Geo Holdout Tests.
The programme version is the one this note is about.
Sizing it
The cost is real: the holdout misses every improvement, so it’s a deliberate revenue sacrifice.
5% holdout, 500,000 visitors/year = 25,000 visitors untreated
if the programme is genuinely worth +3% conversion
forgone 25,000 × 2.5% × 3% × £15.00 contribution ≈ £280
A few hundred pounds a year to know whether your programme works. The cost is trivially small relative to what’s being validated, and the argument against holdouts is almost always framed as cost when it’s really discomfort about the answer.
But statistical power is the binding constraint, not cost. A 5% holdout against 95% treated gives far less power than the split implies, because power depends on the smaller arm:
5% / 95% effective sample ≈ 4× the small arm
50% / 50% effective sample = the full sample
Which means: size the holdout to detect the cumulative effect you’d expect over the period, not to minimise cost. For a programme worth a few percent over a year, 5–10% of traffic held out over 12 months is usually enough — but run the calculation rather than assuming — Sample Size Calculation, Statistical Power.
Implementation
- Assign at the same unit as your tests, and persist it — usually customer ID, since the horizon exceeds any cookie’s life — Randomisation Unit, Assignment and Bucketing
- Assign it first, above everything else. The holdout flag gates entry to every other experiment, so a held-out user is never eligible for any test
- Hold out from shipped winners too, not just live tests. This is the part that’s easy to get wrong and that invalidates the whole thing — once a winner is rolled out to 100%, the holdout must stay on the original
- Log the assignment, so post-hoc analysis can verify the split held — Sample Ratio Mismatch
- Refresh on a schedule, typically annually. Reset the holdout, take the reading, start again. A holdout that runs for three years is measuring an increasingly unrepresentative experience
- Exclude it from personalisation and lifecycle programmes too, or decide explicitly that it’s testing-only. Both are valid; an undocumented mix isn’t
The expensive failure
Keeping a group on a genuinely worse experience for a year has a cost beyond the forgone conversion — those customers may churn, and if the holdout is assigned at customer level, that churn is concentrated in one arm.
Practical mitigations: exclude the holdout from anything that fixes a defect rather than tests an idea (bug fixes, accessibility, performance regressions, legal changes), and set a rule for pulling it early if a guardrail moves — Guardrail Metrics.
Reading it
The comparison is holdout versus everyone else, over the full period, on your primary business metrics.
Expect the holdout result to be smaller than the sum of the wins, and treat that as the correction working rather than as a failure. A programme that claims +18% and holds out at +4% is a functioning programme with an inflated reporting habit — Reading a Test Result, Experiment Archive.
If the holdout shows no difference at all, that’s the finding the holdout exists to produce, and it’s worth far more than another quarter of tests.