Tags: experimentation concept
Assignment and Bucketing
Date: 2026-08-16
Hash the identifier, take the remainder, read off the variant. Deterministic rather than random-at-call-time, so the same user gets the same experience forever without anything being stored.
What it is
Assignment is deciding which variant a unit receives. Bucketing is the mechanism: mapping an identifier onto a position in a fixed range, then splitting that range between variants.
The critical property is that it’s deterministic — the same inputs always produce the same output. No lookup, no database, no state.
The mechanism
user_id "u_84213"
salt "checkout-cta-2026-08" ← experiment key
─────────────────────
hash("u_84213:checkout-cta-2026-08") → 0x7f3a91c4...
take modulo 100 → 42
─────────────────────
buckets 0–49 → control
50–99 → variant
─────────────────────
42 → CONTROL, every time, on every device that knows u_84213
Three things this buys:
- Stickiness for free. Recompute on every page load and you get the same answer, so no assignment needs storing and nothing expires
- Consistency across services. The web front end, the mobile app and a batch email job all compute the same bucket from the same user ID
- Reproducibility. You can rebuild who was in which arm months later from the identifier alone, which matters for reanalysis
Why the salt matters
Hashing the user ID alone would put the same users in the same bucket position in every experiment. User 84213 would land in control for all of them, and the same cohort would be systematically over-represented in one arm across the whole programme.
Salting with the experiment key re-shuffles everyone for each test. It’s a one-line detail that prevents a subtle, programme-wide correlation — and it’s also what makes concurrent tests roughly independent, limiting Interaction Effects.
Requirements on the hash
- Uniform — output evenly spread across buckets. A poor function correlating with identifier structure (sequential IDs, timestamps embedded in IDs) produces arms that differ systematically. This is the deep version of the failure Sample Ratio Mismatch detects
- Stable across languages and versions. If the JavaScript SDK and the server compute different hashes for the same user, a user flips arms depending on where the decision was made. Use the same algorithm everywhere and pin it — MurmurHash and MD5 are common choices, both fine here since this isn’t a security context
- Fast. It runs on every request
Ramps and unequal splits
The bucket range makes traffic allocation trivial and safe:
1% ramp bucket 0 → variant buckets 1–99 → control
5% ramp buckets 0–4 → variant buckets 5–99 → control
50/50 buckets 0–49 → variant buckets 50–99 → control
Expand the range upward and nobody moves arms. A user in bucket 3 who was in the variant at 5% is still in it at 50%. Users only ever join the variant, never leave it — which preserves the consistent experience through a ramp. Shrinking the range does the opposite and should be treated as ending the test.
See Rollouts as Experiments for when that ramp is measurable and when it isn’t.
Failure modes
- Assigning before the identifier exists. A user assigned on a temporary ID that’s later replaced by a real one flips arms mid-visit. Assign as late as possible, on the durable identifier
- The anonymous-to-logged-in switch. Bucketing on cookie ID pre-login and user ID post-login reassigns at the moment of login — which is usually the moment before conversion. Pick one identifier for the whole test, and prefer the cookie unless the test is only for logged-in users. See Randomisation Unit and Identity Stitching
- Falling back to control on failure. When the assignment service times out, defaulting to control loads that arm with slow connections and old devices. Prefer a deterministic local fallback using the same hash
- Assignment without exposure. Bucketing every visitor while only some ever reach the tested page dilutes the effect towards nothing — record exposure separately, see Experiment Assignment Tracking
- Reusing an experiment key. A new test with an old salt inherits the old test’s bucketing, so anyone in the previous variant is in this one too. Include a date or version in the key