Tags: experimentation concept
Experimentation Maturity
Date: 2026-08-17
The stages an organisation moves through, and — more usefully — the specific bottleneck at each one. The stages are descriptive rather than prescriptive; what makes the model worth having is that the constraint changes at every step, so the thing that got you here reliably fails to get you further.
Experimentation maturity describes how far an organisation has moved from occasional ad hoc tests towards experimentation as the default way decisions get made.
The stages, and what actually blocks each
1 AD HOC a few tests a quarter, run by one enthusiast
bottleneck: TRUST. nobody believes the numbers, and they're right —
no A/A validation, no SRM check, no assignment tracking
the move: prove the machinery. A-A Tests, then one clean result
2 ESTABLISHED a regular cadence, a tool, someone's job
bottleneck: THROUGHPUT. everything queues — design, dev, sign-off
the move: concurrency, templates, standing decision rules
— Experimentation Velocity
3 SCALED dozens of concurrent tests, several teams
bottleneck: QUALITY CONTROL. volume outruns rigour. peeking,
post-hoc segmentation and uncorrected multiplicity
start manufacturing findings at scale
the move: automated guardrails, enforced pre-registration,
platform-level stopping rules — Pre-Registration
4 EMBEDDED testing is how decisions are made, not a stage in a process
bottleneck: LEARNING. results accumulate, knowledge doesn't
the move: belief registers, deliberate replication, holdouts
— Institutional Learning
The stage-3 move is the one with teeth: the analysis plan has to be fixed and enforced by the platform rather than by an analyst’s memory — Pre-Registration.
Each bottleneck is invisible from the stage below. At stage 2, quality control looks like bureaucracy; at stage 3 it’s the thing preventing the programme from producing confident nonsense at volume.
The stage-3 trap, which is the interesting one
The dangerous transition is 2 → 3, because velocity and rigour pull against each other and velocity is the one that’s measured.
stage 2 behaviour that stops working at stage 3
analyst manually checks each result → 60 tests/month, nobody checks
one person guards the standards → they become the bottleneck,
then they get overruled
"we'll correct for multiplicity if it → nobody remembers, every test
comes up" segments by device
readouts in a meeting → too many meetings; results get
self-served and misread
The resolution is to move rigour into the platform rather than into people. Automatic SRM checks that flag a test before anyone reads it, primary metrics locked at launch, the analysis running against the pre-registered plan by default, segment results shown with corrections applied. Discipline that depends on someone remembering does not survive volume — Sample Ratio Mismatch, The Multiple Comparisons Problem.
What doesn’t indicate maturity
The signals that get mistaken for progress:
- Test count alone. Sixty underpowered tests is worse than twelve powered ones — Statistical Power
- A high win rate. Usually evidence of a measurement problem or of testing only safe changes — Win Rate and Expected Value
- The tooling. An expensive platform used with no stopping rules is stage 1 with a better dashboard
- Executive enthusiasm. Genuinely helpful, and it’s the thing that most often produces stage-3 quality failures — pressure for wins is pressure to find them
- “We test everything.” Often means “we ship everything and call the rollout a test” — Rollouts as Experiments
What does
- A test has stopped something senior. The clearest single indicator: the programme has authority over a decision, not just a role in it — The HiPPO Problem
- Nulls are reported without apology, and inconclusive results get discussed
- The stopping rule holds when the result is nearly significant on a Friday
- Results are challenged on methodology by people other than the analyst
- Someone can tell you the programme’s win rate from records rather than from memory
- A belief has been overturned by a later test, and the change was recorded rather than argued away
Moving between stages
The transitions are organisational far more than technical:
- 1 → 2 needs one credible, consequential result. Pick a test that matters, run it properly, and let the machinery be scrutinised. Trust is built once and lost quickly
- 2 → 3 needs the queues dismantled — standing approvals, concurrency, templates. This is mostly a permission problem dressed as a capacity problem
- 3 → 4 needs someone whose job is synthesis rather than execution, and it needs leadership to ask “what did we learn” as often as “what did we win”
Skipping stages doesn’t work, and the common attempt is buying a platform to jump from 1 to 3. The tool arrives, nobody trusts the numbers, and the programme lands back at stage 1 having spent a budget.
Regression
Maturity is not monotonic, and it decays quietly:
- The person who held the standards leaves, and nothing was written down
- A replatform breaks assignment or tracking, and nobody re-validates — Replatforming
- A reorganisation splits the programme across teams with no shared archive
- A bad quarter produces pressure for wins, which produces peeking
Re-run an A test periodically, particularly after any platform or tracking change. It’s cheap, and it’s the only routine check that catches silent regression in the machinery everything else depends on.
Where it interacts
- Experimentation Velocity — the stage-2 bottleneck, and the metric most likely to be over-optimised at stage 3
- Institutional Learning — the stage-4 destination
- Experiment Archive — the artefact that makes stage 4 possible and survives staff turnover
- Ethics of Experimentation — a governance question that appears at stage 3, when volume means nobody is individually reviewing what’s being tested on people