Tags: experimentation concept
Test Duration
Date: 2026-08-16
Sample size tells you how many people. Duration tells you which people — and hitting your sample in three days means you sampled three days of behaviour, not your customers.
What it is
Test duration is how long a test runs, decided jointly by two independent constraints: the sample size the design needs, and the calendar time needed for the sample to be representative.
Whichever is longer wins. They are not interchangeable, and treating duration as purely a function of sample size is the most common planning error.
Constraint one: sample
From Sample Size Calculation — 53,000 per arm at 3% baseline and a 10% MDE, at 12,000 visitors a week:
106,000 total ÷ 12,000 per week = 8.8 weeks
Constraint two: cycles
Behaviour is periodic, and each period must be represented whole.
- Weekly. Weekend traffic converts differently from weekday — different intent, devices, time available. Always run in whole weeks. A test ending on a Thursday has one more weekday than weekend versus one starting on a Monday, and that imbalance can exceed the effect you’re measuring
- Payday. Monthly cycles are pronounced in retail and especially in higher-value or discretionary categories
- Campaign. If a promotion runs mid-test, both arms get it, but it changes the population — the traffic during a sale is a different mix, and the effect measured on it may not generalise
- Category seasonality. Half-term, Christmas, seasonal demand
The floor is one full week, and two is safer. Below a week you have not measured a representative population, whatever the sample size says.
sample says 2.1 weeks
cycles say 2 weeks minimum
→ run 2 weeks (whole weeks, same weekday start and end)
sample says 8.8 weeks
cycles say 2 weeks minimum
→ run 9 weeks
The upper bound
Longer isn’t safer indefinitely — several things degrade with time.
- Cookie churn. Assignment is stored client-side and decays. Over months, users are reassigned and arms contaminate each other — Browser Privacy Restrictions, Assignment and Bucketing
- The site changes underneath. Other releases land, campaigns shift, prices move. A quarter-long test is measuring a moving target
- Novelty and Primacy Effects distort the early period and fade, so a short test and a long one can genuinely disagree
- Opportunity cost. A nine-week test occupies a slot. Four two-week tests may be worth more than one nine-week test even if each is less certain — Experimentation Velocity
Rough working limit: beyond six to eight weeks, question whether the test should exist rather than extending it. That usually means raising the MDE, widening the exposure, or accepting the decision is a judgement call — Minimum Detectable Effect.
The one thing you cannot do
Extend a test because it hasn’t reached significance yet. That converts your fixed horizon into a stopping rule based on the outcome, which is Peeking with a longer interval, and it inflates the false positive rate exactly the same way.
The legitimate versions:
- Extend for a pre-registered reason — traffic came in below forecast, so the planned sample takes longer. Decided on traffic, not on the result
- Use a design built for flexible stopping — Sequential Testing or Always-Valid Inference, chosen before launch
- Accept the result and record what MDE the test could actually detect — Inconclusive Results
In plain terms: deciding to run longer because of what the numbers currently say is the same mistake as stopping early, in the opposite direction. Both let the data choose the endpoint.
In practice
- Convert sample size to weeks immediately. Weeks is what gets vetoed and what people can reason about
- Start and end on the same weekday, and write both dates into the plan — Pre-Registration
- Check the calendar before launching. A test spanning Black Friday is measuring Black Friday
- Record actual duration in the archive, not planned. They differ more often than anyone admits