Tags: experimentation concept
Experiment Archive
Date: 2026-08-16
The searchable record of every test run, including the losers. It’s the asset that outlives every individual result — a programme’s accumulated knowledge is worth far more than any one win, and it’s the thing that’s lost first when people leave.
What it is
An experiment archive is a durable, searchable record of what was tested, what was predicted, what happened and what was decided.
The value is not administrative. Two specific returns:
- Nobody re-runs a test that already failed. Without an archive this happens constantly, especially across team changes — the same idea is re-proposed with the same confidence every eighteen months
- The losers are the knowledge. Wins get shipped and become invisible; losses only exist if they’re written down. A programme’s real learning is disproportionately in the tests that didn’t work — Inconclusive Results
What a record has to contain
Anything less and it can’t be re-read years later by someone who wasn’t there:
| Field | Why |
|---|---|
| Hypothesis | What was believed, in the original words — Hypothesis Design |
| Screenshot or diff of each variant | The one field always missing. Prose descriptions of a UI change are useless within a year |
| Where it ran | Pages, templates, devices, audience, geography |
| Primary metric and MDE | The minimum detectable effect the test was sized for — declared before, not chosen after — Pre-Registration, Minimum Detectable Effect |
| Randomisation unit and split | Randomisation Unit |
| Dates and duration | Including whether it covered whole business cycles — Test Duration |
| Sample size per arm | Makes power reconstructable |
| Result with interval | Not just a verdict — Confidence Intervals |
| Guardrails | And whether any moved — Guardrail Metrics |
| The decision, and why | Ship, don’t ship, iterate — including decisions taken against the data |
| Post-rollout outcome | Added later — Post-Test Validation |
The decision field is the one people skip and the one that’s read most. A test can be inconclusive and still lead to shipping, for good reasons — the archive is worthless if it records the statistics and not the reasoning.
Make it findable, not just stored
An archive nobody searches is a folder. What makes it usable:
- Tag by page and by element, not only by date. The real query is “has anyone touched the delivery messaging on the basket?”
- Tag by pattern. Urgency, social proof, form length, image versus video. The transferable finding is at the pattern level, not the page level
- Keep it where people already are. A dedicated tool nobody opens loses to a well-structured shared document. The tool matters far less than the habit
- One record per test, one owner, written while it’s fresh. Retrospective write-ups are thin write-ups
Reviewing the corpus
The archive earns its keep when it’s read in aggregate rather than looked up:
- Win rate. A programme where most tests win is testing changes too safe to be informative, or reading results too generously. A very low rate means the hypotheses aren’t grounded. Both are diagnostics about the pipeline, not the site
- Average effect size of winners, against post-rollout outcomes. That ratio is your calibration factor — Winner’s Curse
- Patterns that repeat. Three tests of the same pattern with the same direction is worth more than one significant result, and is the closest thing to internal meta-analysis you’ll get
- Where you’ve never tested. The archive shows the shape of what’s been examined, and the empty regions are usually more interesting than the crowded ones
The failure mode
Archives die by being a reporting obligation rather than a working tool. The record is filled in for governance, nobody reads it, quality falls, and it stops being worth reading — which was the prediction all along.
The fix is a habit, not a template: check the archive before writing a hypothesis, every time. That’s the only thing that makes the writing worth doing, and it’s the step that gets skipped.