Tags: experimentation concept
Experimentation Velocity
Date: 2026-08-17
Tests completed per period. It’s the lever that dominates every other decision in a programme — because effect sizes are unpredictable and win rates are low, cumulative gain is driven far more by how many attempts you get than by how good each one is.
Experimentation velocity is the number of tests a programme completes per period.
Why throughput beats quality
If a quarter of tests win, and the size of a win isn’t reliably foreseeable, then a programme’s output is close to a sampling process. Sample more.
programme A: 12 tests/year, meticulously chosen
win rate 33% → 4 winners × avg +2.0% = +8.0% compounded-ish
programme B: 48 tests/year, less deliberation
win rate 25% → 12 winners × avg +1.6% = +19.2%
B has a WORSE win rate and SMALLER average wins,
and produces more than twice the gain
Four times the tests at two-thirds the quality still wins comfortably. This is the argument against spending weeks ranking a backlog, and the reason Test Prioritisation matters less than it feels like it should.
The caveat that keeps it honest: this holds while tests are properly powered. Doubling throughput by halving durations produces underpowered tests, which don’t return small wins — they return noise, and shipping noise is negative value — Statistical Power.
Where the time actually goes
Measure the cycle, not the test. The running is rarely the bottleneck.
idea logged day 0
↓ waiting for prioritisation meeting 9 days ← queue
prioritised day 9
↓ waiting for design 14 days ← queue
design done day 23
↓ waiting for dev capacity 11 days ← queue
built day 34
↓ QA and stakeholder sign-off 6 days ← queue
launched day 40
↓ RUNNING 14 days ← the only
analysed day 54 irreducible bit
↓ waiting for the readout meeting 8 days ← queue
decided day 62
62 days total. 14 running. 48 waiting.
Roughly three-quarters of the elapsed time is queueing, and none of it is statistics. Attacking the 14 days is the instinct and it’s the smallest and most dangerous target; attacking the 48 is where the throughput is.
The levers, in order of return
- Run tests concurrently. The single biggest one. Independent salted randomisation makes this safe, and mutual exclusion groups handle the rare genuine conflicts — Interaction Effects. Serialising tests “to be safe” halves throughput to avoid a problem that mostly doesn’t exist
- Kill the approval queues. A standing decision rule — pre-agreed guardrails, pre-agreed stopping rules — removes the need for a meeting per test. Sign-off should be required for the policy, not each instance
- Templates and a component library. Most tests are variations on a dozen shapes. Building each from scratch is where developer days disappear — Design Systems
- A standing analysis template. Automated readouts against the pre-registered plan turn a two-day analysis into an hour, and remove the incentive to interpret creatively
- Reduce the QA surface with a checklist rather than a bespoke investigation — Experiment QA
- Fix the traffic constraint by testing higher up the funnel or on higher-traffic templates, rather than by shortening durations
The constraint that caps it
Velocity is bounded by traffic, and this is a hard ceiling rather than a process problem.
site traffic 400,000 visitors/week
sample per test 210,000 per arm × 2 = 420,000
tests running concurrently, each needing full traffic: ~1 per week... in theory
BUT concurrent tests can share the same traffic
(each user is in every test, independently randomised)
so the real constraint is: how many tests can run for their
required duration simultaneously — which is many, not one
This is the most valuable thing to understand about capacity. Concurrent tests don’t divide traffic between them — each test sees the whole eligible population, because a user can be in a dozen tests at once with independent assignment. The site that “only has traffic for one test at a time” is usually running one test at a time for organisational reasons and blaming arithmetic.
The genuine limits: tests on the same surface (mutual exclusion), tests on low-traffic pages, and the number of distinct pages worth testing at all.
Measuring it
Track these, and nothing else:
- Tests launched per month and tests concluded per month — the second is the real one, and a gap between them means tests are being abandoned or left running
- Cycle time from idea to decision, with the queue breakdown above. This is the diagnostic
- Proportion adequately powered. The check that stops velocity being gamed by running junk
- Concurrency — average tests live at once. Usually the number with most headroom
Don’t target win rate. It’s the metric most easily improved by testing only safe, obvious changes, which reduces learning and expected value at once — Win Rate and Expected Value.
Where it interacts
- Experimentation Maturity — velocity is the main thing that changes between stages, and the bottleneck moves as it rises
- Test Prioritisation — matters in inverse proportion to velocity; at high throughput, ordering barely affects the outcome
- Experiment Archive — high velocity without a record produces the same test three times in two years
- Feature Flags and Progressive Delivery — the engineering practices that make launching a test a routine event rather than a release