Tags: experimentation concept

Experimentation Maturity

Date: 2026-08-17


The stages an organisation moves through, and — more usefully — the specific bottleneck at each one. The stages are descriptive rather than prescriptive; what makes the model worth having is that the constraint changes at every step, so the thing that got you here reliably fails to get you further.


Experimentation maturity describes how far an organisation has moved from occasional ad hoc tests towards experimentation as the default way decisions get made.

The stages, and what actually blocks each

1  AD HOC              a few tests a quarter, run by one enthusiast
   bottleneck:         TRUST. nobody believes the numbers, and they're right —
                       no A/A validation, no SRM check, no assignment tracking
   the move:           prove the machinery. A-A Tests, then one clean result

2  ESTABLISHED         a regular cadence, a tool, someone's job
   bottleneck:         THROUGHPUT. everything queues — design, dev, sign-off
   the move:           concurrency, templates, standing decision rules
                       — Experimentation Velocity

3  SCALED              dozens of concurrent tests, several teams
   bottleneck:         QUALITY CONTROL. volume outruns rigour. peeking,
                       post-hoc segmentation and uncorrected multiplicity
                       start manufacturing findings at scale
   the move:           automated guardrails, enforced pre-registration,
                       platform-level stopping rules — Pre-Registration

4  EMBEDDED            testing is how decisions are made, not a stage in a process
   bottleneck:         LEARNING. results accumulate, knowledge doesn't
   the move:           belief registers, deliberate replication, holdouts
                       — Institutional Learning

The stage-3 move is the one with teeth: the analysis plan has to be fixed and enforced by the platform rather than by an analyst’s memory — Pre-Registration.

Each bottleneck is invisible from the stage below. At stage 2, quality control looks like bureaucracy; at stage 3 it’s the thing preventing the programme from producing confident nonsense at volume.

The stage-3 trap, which is the interesting one

The dangerous transition is 2 → 3, because velocity and rigour pull against each other and velocity is the one that’s measured.

stage 2 behaviour that stops working at stage 3

analyst manually checks each result       → 60 tests/month, nobody checks
one person guards the standards           → they become the bottleneck,
                                            then they get overruled
"we'll correct for multiplicity if it     → nobody remembers, every test
 comes up"                                  segments by device
readouts in a meeting                     → too many meetings; results get
                                            self-served and misread

The resolution is to move rigour into the platform rather than into people. Automatic SRM checks that flag a test before anyone reads it, primary metrics locked at launch, the analysis running against the pre-registered plan by default, segment results shown with corrections applied. Discipline that depends on someone remembering does not survive volume — Sample Ratio Mismatch, The Multiple Comparisons Problem.

What doesn’t indicate maturity

The signals that get mistaken for progress:

  • Test count alone. Sixty underpowered tests is worse than twelve powered ones — Statistical Power
  • A high win rate. Usually evidence of a measurement problem or of testing only safe changes — Win Rate and Expected Value
  • The tooling. An expensive platform used with no stopping rules is stage 1 with a better dashboard
  • Executive enthusiasm. Genuinely helpful, and it’s the thing that most often produces stage-3 quality failures — pressure for wins is pressure to find them
  • “We test everything.” Often means “we ship everything and call the rollout a test” — Rollouts as Experiments

What does

  • A test has stopped something senior. The clearest single indicator: the programme has authority over a decision, not just a role in it — The HiPPO Problem
  • Nulls are reported without apology, and inconclusive results get discussed
  • The stopping rule holds when the result is nearly significant on a Friday
  • Results are challenged on methodology by people other than the analyst
  • Someone can tell you the programme’s win rate from records rather than from memory
  • A belief has been overturned by a later test, and the change was recorded rather than argued away

Moving between stages

The transitions are organisational far more than technical:

  • 1 → 2 needs one credible, consequential result. Pick a test that matters, run it properly, and let the machinery be scrutinised. Trust is built once and lost quickly
  • 2 → 3 needs the queues dismantled — standing approvals, concurrency, templates. This is mostly a permission problem dressed as a capacity problem
  • 3 → 4 needs someone whose job is synthesis rather than execution, and it needs leadership to ask “what did we learn” as often as “what did we win”

Skipping stages doesn’t work, and the common attempt is buying a platform to jump from 1 to 3. The tool arrives, nobody trusts the numbers, and the programme lands back at stage 1 having spent a budget.

Regression

Maturity is not monotonic, and it decays quietly:

  • The person who held the standards leaves, and nothing was written down
  • A replatform breaks assignment or tracking, and nobody re-validates — Replatforming
  • A reorganisation splits the programme across teams with no shared archive
  • A bad quarter produces pressure for wins, which produces peeking

Re-run an A test periodically, particularly after any platform or tracking change. It’s cheap, and it’s the only routine check that catches silent regression in the machinery everything else depends on.

Where it interacts