Tags: experimentation concept
Traffic Allocation
Date: 2026-08-17
How the eligible traffic is divided between arms, and how much of it enters the test at all. Uneven splits feel safer and cost power; the arithmetic of exactly how much is unintuitive, and it’s the reason 90/10 “cautious” tests so often finish with nothing to show.
Traffic allocation is the share of eligible traffic entering a test, and how that share is split between the arms.
Even splits are optimal, and here’s the size of it
For a fixed total sample, power is maximised by splitting evenly. The reason is that the standard error of a difference depends on both arms, so starving one arm inflates the whole interval — the small arm becomes the binding constraint.
The relative sample size needed to hold power constant, versus a 50/50 split:
split effective sample multiplier to keep the same power you need…
50/50 1.00× 100,000 users
60/40 1.04× 104,000
70/30 1.19× 119,000
80/20 1.56× 156,000
90/10 2.78× 278,000
95/5 5.26× 526,000
In plain terms: running 90/10 instead of 50/50 means you need nearly three times as many visitors to learn the same thing. A “cautious” split doesn’t reduce risk — it extends exposure to a possibly-worse variant for two to three times as long, on more users in total.
Working the 90/10 figure, so it doesn’t appear from nowhere. The variance of the difference between two proportions is proportional to 1/n₁ + 1/n₂. With a total of N split by proportions p and 1−p, that’s:
50/50 1/(0.5N) + 1/(0.5N) = 2/N + 2/N = 4.00/N
90/10 1/(0.9N) + 1/(0.1N) = 1.11/N + 10/N = 11.11/N
ratio 11.11 ÷ 4.00 = 2.78
So the variance is 2.78× larger at the same N — and since required sample scales linearly with variance, you need 2.78× the traffic. Note which term dominates: the 10/N from the small arm. The small arm is doing all the damage, which is why 60/40 barely matters and 95/5 is ruinous.
Exposure: what fraction enters the test at all
A separate decision from the split, and routinely confused with it.
ALLOCATION SPLIT
what % of eligible traffic how that traffic divides
enters the experiment between arms
10% into the test, 50/50 within it → 5% control, 5% variant
100% into the test, 90/10 → 90% control, 10% variant
both give the variant 10% of nothing much,
but the first has 5% in control and the second has 90%
Ramping allocation while holding the split even is the good pattern: 5% of traffic into the test at 50/50, then 20%, then 100%, checking for breakage between steps. You get the safety of limited exposure without paying the power penalty above — Progressive Delivery.
The ramp, and the trap inside it
Changing allocation mid-test is safe. Changing the split mid-test is not.
day 1–3 90/10 control 90,000 variant 10,000
day 4–14 50/50 control 200,000 variant 200,000
pooled control 290,000 / variant 210,000
Simpson’s paradox is now live. The two periods have different traffic mixes — different weekdays, different campaigns — and the arms are weighted differently across them, so the pooled comparison can point the opposite way from both periods analysed separately — Simpson’s Paradox.
Two safe options, both fine:
- Discard the ramp period and analyse only from the point the split stabilised. Simplest, and the usual choice
- Analyse each period separately and combine with fixed weights — correct, more work, and rarely worth it for a two-day ramp
Never pool across a split change without doing one of them. And note that users assigned during the ramp stay assigned, so a naive “analyse from day 4” that includes day-1 users’ later behaviour is neither of the above.
Implementing it
Allocation and split are both just ranges over the same hash — which is why the arithmetic is easy to get subtly wrong.
// one hash, two decisions, in this order
function assign(userId, experiment) {
const h = hash(`${experiment.id}:${userId}`); // stable per user per experiment
const bucket = h % 10000; // 0–9999, finer than percent
// 1. is this user in the experiment at all?
if (bucket >= experiment.allocation * 10000) return null; // not exposed
// 2. which arm? rescale within the allocated range, don't reuse the raw bucket
const within = bucket / (experiment.allocation * 10000); // 0–1
return within < experiment.split ? 'control' : 'variant';
}Salting the hash with the experiment ID is what makes concurrent tests independent — without it, the same users land in the same relative position in every test and the arms correlate across experiments — Assignment and Bucketing, Interaction Effects.
Increasing allocation must not reassign anyone. Because the check is bucket < allocation, raising allocation from 10% to 50% adds new buckets and leaves buckets 0–999 exactly where they were. Reordering or re-salting on a ramp reshuffles existing users, which corrupts the test silently — and it’s the commonest implementation bug in this area.
Unequal splits that are justified
- A genuinely risky variant where the downside is large and asymmetric — and accept that you’re buying caution with power, deliberately
- More than two arms. Four arms at 25% each is even and correct; the cost is that each pairwise comparison has a quarter of the traffic, and you now have a multiplicity problem — The Multiple Comparisons Problem
- A shared control across several tests, where one large control serves multiple variants. Legitimate, and the multiplicity correction is not optional
- Bandits, which reallocate continuously by design and answer a different question — Multi-Armed Bandits
Where it interacts
- Statistical Power and Sample Size Calculation — the multipliers above go straight into the sample size, and a calculator assuming 50/50 will understate what an uneven test needs
- Test Duration — allocation is the lever that trades exposure against how long you wait
- Sample Ratio Mismatch — the observed split not matching the intended one is the single most important pre-analysis check, and ramping makes the expected ratio a moving target you have to compute rather than assume