Tags: experimentation concept

Experimentation Velocity

Date: 2026-08-17


Tests completed per period. It’s the lever that dominates every other decision in a programme — because effect sizes are unpredictable and win rates are low, cumulative gain is driven far more by how many attempts you get than by how good each one is.


Experimentation velocity is the number of tests a programme completes per period.

Why throughput beats quality

If a quarter of tests win, and the size of a win isn’t reliably foreseeable, then a programme’s output is close to a sampling process. Sample more.

programme A: 12 tests/year, meticulously chosen
  win rate 33%   →   4 winners × avg +2.0%   =   +8.0% compounded-ish

programme B: 48 tests/year, less deliberation
  win rate 25%   →  12 winners × avg +1.6%   =  +19.2%

B has a WORSE win rate and SMALLER average wins,
and produces more than twice the gain

Four times the tests at two-thirds the quality still wins comfortably. This is the argument against spending weeks ranking a backlog, and the reason Test Prioritisation matters less than it feels like it should.

The caveat that keeps it honest: this holds while tests are properly powered. Doubling throughput by halving durations produces underpowered tests, which don’t return small wins — they return noise, and shipping noise is negative value — Statistical Power.

Where the time actually goes

Measure the cycle, not the test. The running is rarely the bottleneck.

idea logged                    day 0
  ↓  waiting for prioritisation meeting              9 days   ← queue
prioritised                    day 9
  ↓  waiting for design                             14 days   ← queue
design done                    day 23
  ↓  waiting for dev capacity                       11 days   ← queue
built                          day 34
  ↓  QA and stakeholder sign-off                     6 days   ← queue
launched                       day 40
  ↓  RUNNING                                        14 days   ← the only
analysed                       day 54                            irreducible bit
  ↓  waiting for the readout meeting                 8 days   ← queue
decided                        day 62

62 days total.  14 running.  48 waiting.

Roughly three-quarters of the elapsed time is queueing, and none of it is statistics. Attacking the 14 days is the instinct and it’s the smallest and most dangerous target; attacking the 48 is where the throughput is.

The levers, in order of return

  • Run tests concurrently. The single biggest one. Independent salted randomisation makes this safe, and mutual exclusion groups handle the rare genuine conflicts — Interaction Effects. Serialising tests “to be safe” halves throughput to avoid a problem that mostly doesn’t exist
  • Kill the approval queues. A standing decision rule — pre-agreed guardrails, pre-agreed stopping rules — removes the need for a meeting per test. Sign-off should be required for the policy, not each instance
  • Templates and a component library. Most tests are variations on a dozen shapes. Building each from scratch is where developer days disappear — Design Systems
  • A standing analysis template. Automated readouts against the pre-registered plan turn a two-day analysis into an hour, and remove the incentive to interpret creatively
  • Reduce the QA surface with a checklist rather than a bespoke investigation — Experiment QA
  • Fix the traffic constraint by testing higher up the funnel or on higher-traffic templates, rather than by shortening durations

The constraint that caps it

Velocity is bounded by traffic, and this is a hard ceiling rather than a process problem.

site traffic            400,000 visitors/week
sample per test          210,000 per arm × 2  =  420,000
tests running concurrently, each needing full traffic:  ~1 per week... in theory

BUT concurrent tests can share the same traffic
  (each user is in every test, independently randomised)

so the real constraint is:  how many tests can run for their
required duration simultaneously  —  which is many, not one

This is the most valuable thing to understand about capacity. Concurrent tests don’t divide traffic between them — each test sees the whole eligible population, because a user can be in a dozen tests at once with independent assignment. The site that “only has traffic for one test at a time” is usually running one test at a time for organisational reasons and blaming arithmetic.

The genuine limits: tests on the same surface (mutual exclusion), tests on low-traffic pages, and the number of distinct pages worth testing at all.

Measuring it

Track these, and nothing else:

  • Tests launched per month and tests concluded per month — the second is the real one, and a gap between them means tests are being abandoned or left running
  • Cycle time from idea to decision, with the queue breakdown above. This is the diagnostic
  • Proportion adequately powered. The check that stops velocity being gamed by running junk
  • Concurrency — average tests live at once. Usually the number with most headroom

Don’t target win rate. It’s the metric most easily improved by testing only safe, obvious changes, which reduces learning and expected value at once — Win Rate and Expected Value.

Where it interacts

  • Experimentation Maturity — velocity is the main thing that changes between stages, and the bottleneck moves as it rises
  • Test Prioritisation — matters in inverse proportion to velocity; at high throughput, ordering barely affects the outcome
  • Experiment Archive — high velocity without a record produces the same test three times in two years
  • Feature Flags and Progressive Delivery — the engineering practices that make launching a test a routine event rather than a release