Tags: experimentation statistics concept
Multi-Armed Bandits
Date: 2026-08-17
An allocation strategy that shifts traffic towards whichever arm is winning while the test is still running. It minimises the cost of showing people the losing option; it pays for that by producing a worse estimate of how much better the winner actually is — so the question it answers is “which one” rather than “by how much”.
A multi-armed bandit is an allocation algorithm that moves traffic towards the better-performing arm as results accumulate, instead of holding a fixed split until the end.
The trade
The name comes from the slot-machine framing: several levers, unknown payouts, limited pulls. Every pull is either exploration (learning about an arm) or exploitation (taking the payout you believe is best), and you cannot do both with the same pull.
A/B TEST BANDIT
50/50 throughout starts 50/50, drifts towards the winner
day 1 50 / 50
day 4 65 / 35
day 9 88 / 12
day 14 96 / 4
half your traffic sees the loser few see the loser after the first week
for the whole test ← this is the entire benefit
clean estimate of the difference estimate is biased and its interval
with a valid confidence interval is not straightforwardly valid
← this is the entire cost
answers: how much better, and answers: which one, roughly, and
am I sure earns while deciding
Regret is the technical name for what a bandit minimises: the cumulative conversions lost by not having shown the best arm from the start.
Two algorithms, written out
Epsilon-greedy — the simplest thing that works. Explore a fixed fraction of the time, exploit otherwise.
function epsilonGreedy(arms, epsilon = 0.1) {
if (Math.random() < epsilon) {
return arms[Math.floor(Math.random() * arms.length)]; // explore
}
return arms.reduce((best, a) => // exploit
rate(a) > rate(best) ? a : best
);
}
const rate = a => a.trials === 0 ? 0 : a.conversions / a.trials;Crude in a specific way: it explores the clearly-worst arm exactly as often as the nearly-tied one, which wastes most of the exploration budget.
Thompson sampling — the one worth using. Each arm holds a Beta distribution over its true rate; you draw one sample from each and play the highest draw.
function thompson(arms) {
return arms
.map(a => ({
arm: a,
// Beta(successes + 1, failures + 1) — a draw from what we believe
// this arm's true rate might be, given what we've seen so far
draw: sampleBeta(a.conversions + 1, a.trials - a.conversions + 1),
}))
.reduce((best, x) => (x.draw > best.draw ? x : best)).arm;
}Why this is elegant: an arm with few trials has a wide distribution, so it sometimes draws high and gets played — exploration happens automatically, in proportion to genuine uncertainty. An arm that’s clearly worse has a narrow distribution sitting low, so it stops being drawn. No epsilon to tune, and exploration decays on its own as evidence accumulates — Probability Distributions.
Why the estimate goes wrong
The bias is worth understanding rather than just accepting.
arm B has a lucky first day → allocation shifts towards B →
B accumulates most of its data during whatever period followed →
if that period differs from day 1 (weekend, campaign, novelty wearing off),
B's measured rate is a weighted average of the WRONG periods
meanwhile arm A's data is mostly from early on
→ the arms are no longer measured over comparable conditions
This is the same time-confounding as changing a split mid-test, which is exactly what a bandit does continuously by design — Simpson’s Paradox, Traffic Allocation.
Consequences to state plainly:
- The reported lift is unreliable, usually overstated for the arm that won early
- Naive confidence intervals don’t apply. Correcting for adaptive allocation is possible and is not what most vendor implementations do
- Novelty and Primacy Effects are amplified. A variant that wins on novelty gets more traffic, entrenching an effect that would have faded
- Non-stationarity breaks it. If the best arm changes over time — weekday versus weekend, or a seasonal shift — a converged bandit has stopped exploring enough to notice. Discounting old data helps and needs deliberate configuration
When to use one
Good fits, all sharing the property that you don’t need to know by how much:
- Short-lived content with no second chance — a homepage hero for a one-week sale, a subject line, a promotional banner. By the time an A/B test concludes, the campaign is over
- Many arms, low stakes — ten thumbnail images, where running a properly powered ten-arm test is unaffordable and picking a decent one quickly is the whole ask
- Genuinely continuous optimisation — recommendation slots, ranking, ad creative rotation, where the “test” never ends
- Very expensive losers, where the regret is large enough to outweigh a good estimate
Bad fits:
- Anything you need to justify. “The bandit picked it” is not a result you can put in a business case, defend at a review, or add to an archive with a number attached
- Small or slow-accumulating effects. Bandits converge on obvious differences; a real 2% lift takes as long as it would in an A/B test, and now with a corrupted estimate
- Anything with delayed conversion. The algorithm allocates on feedback it hasn’t received yet — a 14-day purchase cycle means the bandit is optimising on a signal that arrives after the decisions were made. This rules out most ecommerce checkout work
- Learning. A bandit tells you which arm, never why, so nothing transfers to the next decision — Institutional Learning
The pragmatic hybrid: run a fixed 50/50 A/B test to a proper conclusion, then use a bandit for ongoing rotation among the survivors. You get the defensible estimate and the reduced regret, in the order that makes each work.
Where it interacts
- Stopping Rules — a bandit replaces the stopping decision with continuous reallocation, which is why it needs a different rule for “we’re done”
- Test Duration — bandits don’t remove the need for full business cycles; they just make the imbalance across those cycles worse
- Personalisation Tests — contextual bandits extend this by choosing per user segment, and inherit every caveat above plus the targeting problems
- Win Rate and Expected Value — the regret calculation is the same arithmetic as the programme-level expected value, applied within one test