Tags: experimentation statistics concept
Switchback Tests
Date: 2026-08-17
Alternating the whole system between control and treatment over time periods, and comparing periods rather than users. It’s the design for when users can’t be independently assigned — because they share a resource, or affect each other — which is exactly when a normal A/B test silently produces a wrong answer.
A switchback test switches the whole system between control and treatment across successive time periods, and compares the periods rather than the users.
The problem it solves
A standard A/B test assumes one user’s treatment doesn’t affect another user’s outcome. That assumption has a name — SUTVA, the stable unit treatment value assumption — and it fails whenever users compete for something shared.
test: a new algorithm that shows "only 2 left in stock" more aggressively
user-level randomisation, 50/50
variant users see urgency → buy faster → stock runs out
↓
control users find the item out of stock → don't buy
measured: variant +12%, control −6%
reality: the variant CANNOT deliver +12% at 100% rollout,
because at 100% there is no control group to steal stock from
The variant’s measured win is partly transferred from the control arm rather than created. Roll it out fully and the gain shrinks or vanishes. The test wasn’t badly run — the design was wrong for the situation.
Where this bites: shared inventory, delivery capacity, marketplaces (a seller boosted for one buyer is demoted for another), customer service queues, pricing where users compare, anything with a fixed pool of supply.
The design
Randomise time periods instead of users. Everyone gets the same experience at any given moment.
period experience orders/hr the unit of analysis is
the PERIOD, not the user
10:00 control 412
11:00 treatment 438
12:00 control 401
13:00 control 395
14:00 treatment 447
15:00 treatment 429
16:00 control 408
…
n = number of periods, not number of users
Randomise the sequence, don’t alternate it. Strict A/B/A/B alternation correlates with any daily rhythm of the same frequency — a system that runs a batch job every second hour would be perfectly confounded. Randomise each period independently, or use a balanced randomised block design so each day gets equal amounts of both.
The cost: your sample size collapses
This is the thing to understand before proposing one.
A/B TEST SWITCHBACK
n = 200,000 users n = 336 periods
(2 weeks × 24 hourly periods)
standard error scales as 1/√200,000 standard error scales as 1/√336
→ intervals roughly 24× wider
for the same underlying variance
In plain terms: going from 200,000 users to 336 periods costs you almost all your statistical power, because the effective sample is the number of periods. Switchbacks detect large effects. A 2% improvement is generally not findable this way; a 15% one is.
Two levers, and both are constrained:
- Shorter periods → more of them → more power. But too short and the carryover problem below dominates
- Longer test → more periods, linearly. A month of hourly periods is 720, which is better and still small
Reducing period-level variance helps more than either. Comparing each period against the same hour on other days — a form of blocking — removes the daily pattern from the noise, and is the standard practical improvement.
Carryover, and the period length trade
The other assumption: an effect must not leak from one period into the next.
period ends 11:00, switches to control
a user who added to basket at 10:52 under treatment
checks out at 11:06 under control
↓
that conversion is attributed to control, caused by treatment
period too SHORT ← carryover contaminates. effects bleed across boundaries
period too LONG ← too few periods. no power
typical resolution: period length ≈ 3–5× the typical time from
exposure to outcome
food delivery (minutes) → 30–60 minute periods
ecommerce checkout (~20 min) → 2–4 hour periods
considered purchase (days) → switchback is the wrong design entirely
Burn-in is the standard mitigation: discard the first few minutes of each period from the analysis, so the measured window contains only outcomes that both started and finished under one condition. It costs sample and it’s usually worth it.
Analysis
- The period is the unit. Compute a metric per period, then compare period-level means. Analysing at user level inside a switchback recreates the dependence you designed around and produces intervals that are far too narrow
- Block on time-of-day and day-of-week. Without it, the daily cycle is most of your variance
- Watch for autocorrelation. Adjacent periods are correlated — busy hours cluster — and standard errors need to account for it or they’ll be overconfident
- Check the balance. Randomisation over a small number of periods can easily land 60/40 by chance, and with n=336 that’s a real imbalance rather than a rounding error
When to use one
Yes: shared inventory or capacity, marketplace matching, delivery and logistics, pricing visible to everyone, service-desk staffing, anything where the treatment changes a shared resource.
No: the effect is small; the outcome takes days; the change is per-user by nature and users don’t interact; you need a per-segment answer.
Alternatives worth considering first: Geo Holdout Tests randomise regions instead of time, which keeps a cleaner comparison when you have enough regions; cluster randomisation assigns whole groups — a warehouse, a store, a city — where the interference is contained inside the cluster.
Where it interacts
- Controlled Experiments — a switchback is still one, with time as the randomisation unit
- Randomisation Unit — this is the case where the honest answer to “what’s the unit” isn’t a user
- Test Duration — duration determines sample directly here, in a way it doesn’t for user-level tests
- Seasonality in Tests — daily and weekly rhythms are the dominant noise source, so blocking on them isn’t optional