Tags: experimentation concept

Traffic Allocation

Date: 2026-08-17


How the eligible traffic is divided between arms, and how much of it enters the test at all. Uneven splits feel safer and cost power; the arithmetic of exactly how much is unintuitive, and it’s the reason 90/10 “cautious” tests so often finish with nothing to show.


Traffic allocation is the share of eligible traffic entering a test, and how that share is split between the arms.

Even splits are optimal, and here’s the size of it

For a fixed total sample, power is maximised by splitting evenly. The reason is that the standard error of a difference depends on both arms, so starving one arm inflates the whole interval — the small arm becomes the binding constraint.

The relative sample size needed to hold power constant, versus a 50/50 split:

split      effective sample multiplier      to keep the same power you need…

50/50            1.00×                      100,000 users
60/40            1.04×                      104,000
70/30            1.19×                      119,000
80/20            1.56×                      156,000
90/10            2.78×                      278,000
95/5             5.26×                      526,000

In plain terms: running 90/10 instead of 50/50 means you need nearly three times as many visitors to learn the same thing. A “cautious” split doesn’t reduce risk — it extends exposure to a possibly-worse variant for two to three times as long, on more users in total.

Working the 90/10 figure, so it doesn’t appear from nowhere. The variance of the difference between two proportions is proportional to 1/n₁ + 1/n₂. With a total of N split by proportions p and 1−p, that’s:

50/50   1/(0.5N) + 1/(0.5N)  =  2/N + 2/N   =  4.00/N
90/10   1/(0.9N) + 1/(0.1N)  =  1.11/N + 10/N =  11.11/N

ratio   11.11 ÷ 4.00  =  2.78

So the variance is 2.78× larger at the same N — and since required sample scales linearly with variance, you need 2.78× the traffic. Note which term dominates: the 10/N from the small arm. The small arm is doing all the damage, which is why 60/40 barely matters and 95/5 is ruinous.

Exposure: what fraction enters the test at all

A separate decision from the split, and routinely confused with it.

ALLOCATION                          SPLIT

what % of eligible traffic          how that traffic divides
enters the experiment               between arms

10% into the test, 50/50 within it  →  5% control, 5% variant
100% into the test, 90/10           →  90% control, 10% variant

both give the variant 10% of nothing much,
but the first has 5% in control and the second has 90%

Ramping allocation while holding the split even is the good pattern: 5% of traffic into the test at 50/50, then 20%, then 100%, checking for breakage between steps. You get the safety of limited exposure without paying the power penalty above — Progressive Delivery.

The ramp, and the trap inside it

Changing allocation mid-test is safe. Changing the split mid-test is not.

day 1–3     90/10     control 90,000   variant 10,000
day 4–14    50/50     control 200,000  variant 200,000

pooled      control 290,000 / variant 210,000

Simpson’s paradox is now live. The two periods have different traffic mixes — different weekdays, different campaigns — and the arms are weighted differently across them, so the pooled comparison can point the opposite way from both periods analysed separately — Simpson’s Paradox.

Two safe options, both fine:

  1. Discard the ramp period and analyse only from the point the split stabilised. Simplest, and the usual choice
  2. Analyse each period separately and combine with fixed weights — correct, more work, and rarely worth it for a two-day ramp

Never pool across a split change without doing one of them. And note that users assigned during the ramp stay assigned, so a naive “analyse from day 4” that includes day-1 users’ later behaviour is neither of the above.

Implementing it

Allocation and split are both just ranges over the same hash — which is why the arithmetic is easy to get subtly wrong.

// one hash, two decisions, in this order
function assign(userId, experiment) {
  const h = hash(`${experiment.id}:${userId}`);   // stable per user per experiment
  const bucket = h % 10000;                       // 0–9999, finer than percent
 
  // 1. is this user in the experiment at all?
  if (bucket >= experiment.allocation * 10000) return null;   // not exposed
 
  // 2. which arm? rescale within the allocated range, don't reuse the raw bucket
  const within = bucket / (experiment.allocation * 10000);    // 0–1
  return within < experiment.split ? 'control' : 'variant';
}

Salting the hash with the experiment ID is what makes concurrent tests independent — without it, the same users land in the same relative position in every test and the arms correlate across experiments — Assignment and Bucketing, Interaction Effects.

Increasing allocation must not reassign anyone. Because the check is bucket < allocation, raising allocation from 10% to 50% adds new buckets and leaves buckets 0–999 exactly where they were. Reordering or re-salting on a ramp reshuffles existing users, which corrupts the test silently — and it’s the commonest implementation bug in this area.

Unequal splits that are justified

  • A genuinely risky variant where the downside is large and asymmetric — and accept that you’re buying caution with power, deliberately
  • More than two arms. Four arms at 25% each is even and correct; the cost is that each pairwise comparison has a quarter of the traffic, and you now have a multiplicity problem — The Multiple Comparisons Problem
  • A shared control across several tests, where one large control serves multiple variants. Legitimate, and the multiplicity correction is not optional
  • Bandits, which reallocate continuously by design and answer a different question — Multi-Armed Bandits

Where it interacts

  • Statistical Power and Sample Size Calculation — the multipliers above go straight into the sample size, and a calculator assuming 50/50 will understate what an uneven test needs
  • Test Duration — allocation is the lever that trades exposure against how long you wait
  • Sample Ratio Mismatch — the observed split not matching the intended one is the single most important pre-analysis check, and ramping makes the expected ratio a moving target you have to compute rather than assume