Tags: statistics experimentation concept

Sequential Testing

Date: 2026-08-16


A design that lets you check a test in flight without inflating the false positive rate. It buys the licence to look by spending power — and you have to buy it before launch, because it can’t be applied retrospectively.


What it is

Sequential testing is any design where the data may be evaluated repeatedly and stopped early, with the error rate controlled across all those looks rather than at one endpoint.

The contrast with a fixed-horizon test:

FIXED HORIZON              SEQUENTIAL

evaluate once, at n        evaluate whenever
α = 5% at that point       α = 5% across ALL looks
looking early breaks it    looking is the design

Why it’s needed

A fixed-horizon test’s 5% error rate assumes one comparison. Each additional look is another chance to cross the threshold by luck — five naive looks push the false positive rate towards 20% — Peeking.

Sequential designs solve this by making the threshold move. Early on, when the sample is small and the estimate wild, the bar for stopping is very high. As data accumulates it relaxes towards the nominal level.

significance threshold required to stop

  strict │●
         │ ╲
         │  ●──●
         │        ╲●──●──●
  0.05   │                  ●──●──●
         └────────────────────────────
          early              planned end

In plain terms: you’re allowed to look whenever you like, but stopping early requires much stronger evidence than stopping at the end. That’s the trade, and it’s an honest one.

The families

Group sequential. A fixed number of pre-planned interim analyses — say at 25%, 50%, 75% and 100% of the planned sample — with the error budget spread across them by a spending function. Well-established, and it requires committing to the look schedule in advance.

Always-valid inference. Confidence sequences that remain valid at every moment, so you can look continuously with no schedule at all. More flexible, more conservative — Always-Valid Inference.

Bandits are related but answer a different question — they allocate traffic towards the winner during the test rather than deciding whether an effect exists — Multi-Armed Bandits.

What it costs

Nothing is free:

  • Power. A sequential design reaching its full sample has slightly less power than a fixed test with the same n, because the error budget was spread across the looks
  • Or equivalently, more traffic to reach the same power
  • Complexity. The analysis is harder to explain and harder to check
  • You must commit in advance. This is the crucial one

Typical cost is modest — often around 10-20% more sample for the full-length case [CHECK: depends on the spending function and the number of looks; verify against your platform’s documentation rather than assuming].

The genuine benefit

Not “look whenever you want”. It’s stopping early when the effect is large.

A change that’s dramatically better or dramatically worse can be identified in a fraction of the planned sample. On a programme running many tests, that returns traffic to the queue and raises throughput — Experimentation Velocity.

The asymmetric case is stronger still: stopping early for harm is always legitimate, in any design, because you’re avoiding damage rather than claiming a discovery. Sequential methods make that formal rather than ad hoc — Guardrail Metrics.

What it does not fix

  • You cannot apply it retrospectively. Peeking at a fixed-horizon test and then computing a sequential p-value is not valid — the design has to be chosen before launch
  • It doesn’t fix multiple metrics. Sequential methods control error across time, not across metrics — The Multiple Comparisons Problem
  • It doesn’t fix underpowering. A sequential test on insufficient traffic is still insufficient
  • It doesn’t fix Sample Ratio Mismatch or any validity problem

When to use it

  • Where stopping early has real value — high test volume, or expensive exposure
  • Where harm detection matters and you need a principled rule
  • Where the organisation will look regardless. This is the honest one: if stakeholders check the dashboard daily whatever you say, a sequential design makes that behaviour valid instead of corrosive

For a modest programme running a test every nine weeks, a fixed horizon with genuine discipline is simpler and costs less power. Sequential earns its keep at volume.