Tags: experimentation concept

Post-Test Validation

Date: 2026-08-16


Checking that the shipped change delivered what the test promised. Almost nobody does it, and it’s the step that catches the two most common expensive outcomes: the winner that was never built properly, and the winner that was never real.


What it is

Post-test validation is the measurement taken after a winning variant is rolled out to everyone, comparing what happened to what the test predicted.

The test measured a variant. Production runs a re-implementation of that variant, at full traffic, for an indefinite period, with the rest of the site changing around it. Those are not the same thing, and the assumption that they are is unexamined nearly everywhere.

The three ways a shipped winner underperforms

1. It wasn’t implemented the same way.

The test ran as a client-side override; production is a proper build. Somewhere in the translation, something differs — a breakpoint, a default, an interaction, an edge case the override never hit because it only ran on one template — Client-Side vs Server-Side Testing.

Check by diffing the rendered output, on the same devices and breakpoints the test covered. This is a five-minute check that catches a surprising share of them.

2. The effect was inflated.

Effects measured on winners are biased upward — you selected on the estimate, so you selected partly on the noise. Expect the true effect to be smaller than the measured one, sometimes much smaller, and most so for tests that were barely significant or underpowered — Winner’s Curse, Statistical Power.

3. It faded.

Novelty wears off; some wins are attention effects with a half-life. A change that worked because it was new stops working once it’s normal — Novelty and Primacy Effects.

How to actually validate

Ranked by strength, and by how much work they are:

MethodStrengthCost
Keep a small holdout on the originalCausal, definitiveSmall, needs planning
Staged rollout with measurementGoodLow if you already ship this way
Before/after on the metricWeak — confounded by everythingNone
Programme holdoutAnswers the aggregate, not this changeOngoing

The single best habit: don’t roll the winner to 100%. Ship to 95%, keep 5% on the original for a defined period, and read it again at 4, 8 and 12 weeks. Costs almost nothing, and it converts every rollout into a durability check — Rollouts as Experiments.

If a residual holdout isn’t possible, staged rollout at least separates implementation failure from effect decay: a step change at the moment of a rollout stage is an implementation signal; a slow drift back to baseline is decay.

Before/after is the weak version

The default — compare the four weeks before to the four weeks after — is confounded by everything that changed at the same time: season, promotions, traffic mix, other releases.

It is not useless. It won’t confirm a modest win, but it will catch a disaster, and that’s a real job. Use it as a smoke alarm, not as a measurement — and only against the same metric definitions the test used — Metric Design.

What to check, and when

day 1      implementation diff — does production match the variant?
day 1      instrumentation — are the events still firing?
week 1     guardrails — errors, latency, returns, support contacts
week 4     primary metric against the test's prediction
week 12    same again, for decay

The day-one instrumentation check matters more than it looks. A rebuild frequently breaks the tracking that measured the original test, which means the validation reads flat because the data stopped, not because the change failed — Experiment Assignment Tracking, Tracking Plans.

Recording the outcome

Write the validated result back against the original test, in the Experiment Archive. A test whose predicted effect never materialised is one of the most valuable records you can hold: it recalibrates every future estimate from the same team, the same tool and the same kind of change.

Keep the prediction and the outcome side by side. Over enough tests you get a calibration factor for your own programme — “our winners deliver about 60% of what they measure” — which is worth more than any individual result, and is the input that makes a business case honest — Reading a Test Result.