Holdout groups
A holdout is a slice of traffic that stays on control after the experiment ships, often permanently. Typically 5-10% of users. The variant is “shipped” to the other 90-95%, but the holdout never gets it, so you can measure the long-tail effect of the change against a clean baseline weeks or months after launch.
This addresses a specific blind spot. A/B tests usually run for 2-4 weeks. A test that wins in that window can underperform over the following 90 days because of novelty effect fading, retention impact, or interaction with seasonality. The holdout is the way to actually measure that.
When holdouts are worth running
Section titled “When holdouts are worth running”- Big changes that might affect retention. Subscription flows, pricing changes, onboarding redesigns. The short-term conversion lift can be misleading if the customers acquired through the variant churn faster.
- Changes with known novelty risk. Anything visually disruptive or behaviourally novel. Tests show inflated initial lift that fades.
- Changes that interact with downstream funnels. A “winning” PDP that brings in lower-LTV buyers via aggressive discounting is a long-run loss. The holdout catches it.
What you measure
Section titled “What you measure”The holdout gives you the variant’s effect on LTV, retention, repeat purchase, and any other metric that needs time to develop. The metric that mattered in the short-term test is usually still measured, but the interesting numbers are the ones that couldn’t show up in the test window.
Comparison is usually done with the same statistical machinery as the original A/B test, just with the variant population vs the holdout population over a longer time frame.
What the holdout isn’t is a retroactive verdict on whether the original test was right. The two measure different time horizons and can honestly disagree - a real short-term lift that decays to nothing over 90 days means both readings were correct about their own window. Treating the holdout as the “true” answer that overrules the test misses that the decay is the finding.
The cost
Section titled “The cost”Holdouts are unprofitable by design. The 10% that doesn’t get the change foregoes whatever uplift the change actually produces. For changes that genuinely lift revenue, the holdout is a small permanent revenue tax. Teams need to be willing to pay it for the measurement value.
That cost is the main reason holdouts are rare outside mature programmes. Most teams take the short-term win and skip the long-run measurement. They learn what the long-run effect was only when they finally turn the change off and discover what the underlying baseline is doing.
Keeping one clean
Section titled “Keeping one clean”A holdout is only worth its cost while it stays uncontaminated, and it degrades quietly.
- Someone ships to it. Usually well-meant - the change has been live for months, the holdout users are “missing out”, so a release quietly goes to 100%. The moment that happens the long-run comparison is gone and there’s no way to reconstruct it.
- Its users get reused. A permanent bucket that nobody remembers exists ends up inside later experiments, and now a single user is in two permanent assignments with unclear precedence. Whoever owns the holdout needs it visible in the same place tests get planned, or it will be forgotten within a quarter.