Tags: statistics experimentation concept

Always-Valid Inference

Date: 2026-08-16


Confidence intervals that stay honest no matter when you look, including continuously. It’s what several modern experimentation platforms run by default — which means their numbers behave differently from a textbook test, and knowing that changes how you read them.


What it is

Always-valid inference produces a confidence sequence — an interval that is simultaneously valid at every point in time, rather than at one pre-specified endpoint.

FIXED-HORIZON CI            CONFIDENCE SEQUENCE

valid at n = 106,000        valid at every n
valid if you look once      valid however often you look
                            no schedule required

The guarantee is stronger than a group sequential design, which requires committing to a look schedule in advance. Here there’s no schedule to commit to.

How it manages that

The interval is wider than a fixed-horizon one at any given sample size — substantially so early on, converging towards it as data accumulates.

interval width

 wide  │●
       │ ╲
       │  ╲●
       │    ╲●─────●
       │            ╲●──────●──────●   confidence sequence
       │  · · · · · · · · · · · · ·    fixed-horizon CI
       └──────────────────────────────
        early                     end

In plain terms: you’re paying for the right to look whenever you want, and the price is a wider interval. That’s an honest trade rather than a free lunch — the width is the cost.

Why platforms use it

Because it matches how people actually behave. Stakeholders check dashboards; telling them not to look has never worked at any organisation. A design that’s valid under continuous monitoring turns a corrosive habit into a supported one.

The consequences worth knowing when you read such a platform’s output:

  • Intervals look wide early. That’s correct, not a bug or an underpowered test
  • “Significance” arrives later than a fixed-horizon test would suggest on the same data
  • You can stop as soon as the interval clears your threshold, legitimately
  • Comparing its numbers to a hand-computed fixed-horizon test will not match, and the platform’s is the right one for how you’re using it

What it costs

  • Wider intervals, which means more traffic for the same precision. The cost is largest early and shrinks
  • Conservatism. Genuinely valid at all times means valid in the worst case, so it errs towards not detecting
  • Harder to explain. “The interval is wide because we’re allowed to look whenever” is a real conversation
  • Not comparable across designs. A meta-analysis mixing always-valid and fixed-horizon results is comparing different guarantees — Experiment Archive

What it doesn’t fix

The same limits as any sequential method:

  • Multiple metrics. It controls error across time, not across metrics — The Multiple Comparisons Problem
  • Post-hoc segmentation — Segmentation (test results)
  • Validity problems. Sample Ratio Mismatch invalidates the test regardless
  • Retrospective application. You cannot peek at a fixed-horizon test and then compute an always-valid interval. The design is a commitment made before launch, not an analysis chosen after

That last point is the one worth holding: it’s not a way to rescue a test you peeked at.

Choosing between the designs

Use
Fixed horizonModest test volume, disciplined team, simplest to explain and to check
Group sequentialPlanned interim looks, want early stopping, can commit to a schedule
Always-validContinuous monitoring is inevitable, or a dashboard is permanently visible to stakeholders

For a programme running a test every nine weeks with a small team, fixed horizon plus discipline is simpler and cheaper. Always-valid earns its cost where the monitoring is going to happen whatever you decide — and being realistic about that is better than a design everyone quietly violates.

See Peeking for what happens without either, and Guide - Statistics for CRO for where this sits in running a test.