Tags: experimentation statistics concept
Interaction Effects
Date: 2026-08-17
Two concurrent tests affecting each other, so that the effect of one depends on which variant of the other a user got. It’s the standard objection to running tests simultaneously, it’s usually raised by people who don’t run many tests, and it’s much rarer than it sounds — but when it’s real it’s invisible unless you go looking.
An interaction effect is when the effect of one test’s variant depends on which variant of another concurrent test the same user received.
Why concurrent tests are usually fine
The reassuring fact is structural rather than optimistic. If both tests randomise independently, each test’s variants are evenly represented within the other test’s arms.
test A: new PDP layout test B: new checkout copy
both 50/50, independently hashed
B control B variant A's totals
A control 25,000 25,000 50,000
A variant 25,000 25,000 50,000
reading test A: control has 50% B-variant users
variant has 50% B-variant users
→ B's effect is balanced across A's arms
→ it adds variance, but does not bias A's estimate
Test B becomes noise in test A rather than bias, which is exactly what randomisation is for — it’s the same mechanism that handles the thousands of other differences between users that you never measured — Why Randomisation Works.
This holds only if the hashes are independent. Salt every experiment’s hash with its own ID, or users land in the same relative bucket in every test and the arms correlate across experiments — at which point the reassurance above evaporates — Assignment and Bucketing, Traffic Allocation.
When it’s genuinely a problem
Interaction matters when the effects are not additive — when A’s effect is different depending on B.
B control B variant
A control 3.0% 3.2% B alone: +0.2pp
A variant 3.4% 3.6% A alone: +0.4pp
A+B expected: +0.6pp → 3.6% ✓
observed 3.6% → NO interaction
versus
B control B variant
A control 3.0% 3.2% A alone: +0.4pp
A variant 3.4% 2.9% B alone: +0.2pp
expected 3.6%, observed 2.9%
→ −0.7pp interaction. REAL
The second table is what a genuine interaction looks like: the joint effect isn’t the sum of the parts. Both tests read positively on their own, and shipping both together makes things worse than shipping neither.
The realistic cases where this happens:
- Both tests change the same element or the same decision. Two tests on the checkout button, one on colour and one on label, producing an invisible or absurd combination
- Both compete for the same attention. A new trust badge and a new delivery banner in the same viewport, each of which works alone by being the only new thing
- One test changes what the other test’s users see upstream. A PDP test that changes which products get added to basket, feeding a basket test
- Both are large. Two subtle copy changes rarely interact; two page redesigns plausibly do
- They’re on the same page and one breaks the other. Not a statistical interaction at all — a rendering conflict, and it’s the most common real version of this in practice
What to do about it
Cheapest to most expensive:
- Mutually exclude tests that touch the same surface. A single exclusion group per page or per journey step, enforced by the platform, is 90% of the answer for 10% of the effort. Anything not in the same group runs concurrently without ceremony
- Check for it after the fact, cheaply. For any two tests that overlapped substantially in time and traffic, cross-tabulate as above. It’s a four-cell table and it takes minutes. Do it for the pairs that plausibly interact rather than all pairs
- QA the combinations. Render all four cells before launch. This catches the layout collisions, which are the common failure — Experiment QA
- Full factorial design — deliberately analysing the interaction term. Correct, and expensive: detecting an interaction reliably needs roughly four times the sample of detecting a main effect of the same size, because you’re estimating a difference of differences. Reserve it for cases where the interaction is the question — Multivariate Tests
- Run them sequentially. The instinct, and usually the wrong call. It halves your throughput to avoid a problem that mostly isn’t there — Experimentation Velocity
The statistical trap in going looking
Testing every pair of concurrent tests for interactions is a large multiplicity problem. With ten concurrent tests there are 45 pairs, and at 95% confidence you’d expect roughly two “significant” interactions from noise alone — The Multiple Comparisons Problem.
So don’t test all pairs. Test the pairs where a mechanism is plausible, decided before looking at the numbers. An interaction with no story attached is almost always noise, and chasing it produces the same false discoveries as post-hoc segmentation — Segmentation (test results).
Where it interacts
- Traffic Allocation — independent salted hashing is the precondition for everything above
- Multivariate Tests — the deliberate version, where the interaction is what you’re paying to measure
- Experimentation Velocity — the reason not to serialise tests by default, and the cost of over-reacting to this risk
- Personalisation Tests — targeting rules layered on top of tests are a systematic source of real interactions, because the targeting isn’t randomised