Tags: experimentation statistics concept

Switchback Tests

Date: 2026-08-17


Alternating the whole system between control and treatment over time periods, and comparing periods rather than users. It’s the design for when users can’t be independently assigned — because they share a resource, or affect each other — which is exactly when a normal A/B test silently produces a wrong answer.


A switchback test switches the whole system between control and treatment across successive time periods, and compares the periods rather than the users.

The problem it solves

A standard A/B test assumes one user’s treatment doesn’t affect another user’s outcome. That assumption has a name — SUTVA, the stable unit treatment value assumption — and it fails whenever users compete for something shared.

test: a new algorithm that shows "only 2 left in stock" more aggressively

user-level randomisation, 50/50

variant users see urgency → buy faster → stock runs out
                                            ↓
control users find the item out of stock → don't buy

measured: variant +12%, control −6%
reality:  the variant CANNOT deliver +12% at 100% rollout,
          because at 100% there is no control group to steal stock from

The variant’s measured win is partly transferred from the control arm rather than created. Roll it out fully and the gain shrinks or vanishes. The test wasn’t badly run — the design was wrong for the situation.

Where this bites: shared inventory, delivery capacity, marketplaces (a seller boosted for one buyer is demoted for another), customer service queues, pricing where users compare, anything with a fixed pool of supply.

The design

Randomise time periods instead of users. Everyone gets the same experience at any given moment.

period    experience     orders/hr    the unit of analysis is
                                       the PERIOD, not the user
10:00     control          412
11:00     treatment        438
12:00     control          401
13:00     control          395
14:00     treatment        447
15:00     treatment        429
16:00     control          408
…

n = number of periods, not number of users

Randomise the sequence, don’t alternate it. Strict A/B/A/B alternation correlates with any daily rhythm of the same frequency — a system that runs a batch job every second hour would be perfectly confounded. Randomise each period independently, or use a balanced randomised block design so each day gets equal amounts of both.

The cost: your sample size collapses

This is the thing to understand before proposing one.

A/B TEST                             SWITCHBACK

n = 200,000 users                    n = 336 periods
                                       (2 weeks × 24 hourly periods)

standard error scales as 1/√200,000  standard error scales as 1/√336
                                     → intervals roughly 24× wider
                                       for the same underlying variance

In plain terms: going from 200,000 users to 336 periods costs you almost all your statistical power, because the effective sample is the number of periods. Switchbacks detect large effects. A 2% improvement is generally not findable this way; a 15% one is.

Two levers, and both are constrained:

  • Shorter periods → more of them → more power. But too short and the carryover problem below dominates
  • Longer test → more periods, linearly. A month of hourly periods is 720, which is better and still small

Reducing period-level variance helps more than either. Comparing each period against the same hour on other days — a form of blocking — removes the daily pattern from the noise, and is the standard practical improvement.

Carryover, and the period length trade

The other assumption: an effect must not leak from one period into the next.

period ends 11:00, switches to control

  a user who added to basket at 10:52 under treatment
  checks out at 11:06 under control
        ↓
  that conversion is attributed to control, caused by treatment
period too SHORT     ← carryover contaminates. effects bleed across boundaries
period too LONG      ← too few periods. no power

typical resolution:  period length ≈ 3–5× the typical time from
                     exposure to outcome

  food delivery (minutes)      → 30–60 minute periods
  ecommerce checkout (~20 min) → 2–4 hour periods
  considered purchase (days)   → switchback is the wrong design entirely

Burn-in is the standard mitigation: discard the first few minutes of each period from the analysis, so the measured window contains only outcomes that both started and finished under one condition. It costs sample and it’s usually worth it.

Analysis

  • The period is the unit. Compute a metric per period, then compare period-level means. Analysing at user level inside a switchback recreates the dependence you designed around and produces intervals that are far too narrow
  • Block on time-of-day and day-of-week. Without it, the daily cycle is most of your variance
  • Watch for autocorrelation. Adjacent periods are correlated — busy hours cluster — and standard errors need to account for it or they’ll be overconfident
  • Check the balance. Randomisation over a small number of periods can easily land 60/40 by chance, and with n=336 that’s a real imbalance rather than a rounding error

When to use one

Yes: shared inventory or capacity, marketplace matching, delivery and logistics, pricing visible to everyone, service-desk staffing, anything where the treatment changes a shared resource.

No: the effect is small; the outcome takes days; the change is per-user by nature and users don’t interact; you need a per-segment answer.

Alternatives worth considering first: Geo Holdout Tests randomise regions instead of time, which keeps a cleaner comparison when you have enough regions; cluster randomisation assigns whole groups — a warehouse, a store, a city — where the interference is contained inside the cluster.

Where it interacts

  • Controlled Experiments — a switchback is still one, with time as the randomisation unit
  • Randomisation Unit — this is the case where the honest answer to “what’s the unit” isn’t a user
  • Test Duration — duration determines sample directly here, in a way it doesn’t for user-level tests
  • Seasonality in Tests — daily and weekly rhythms are the dominant noise source, so blocking on them isn’t optional