Tags: experimentation concept
Seasonality in Tests
Date: 2026-08-17
Running across a peak, a payday or a campaign, and mistaking the period for the treatment. Randomisation protects the comparison — both arms get the same Black Friday — so seasonality rarely biases a result. What it does is change who the result applies to, which is a generalisation problem rather than a validity one.
Seasonality in tests is the risk that the period a test runs in — a sale, payday, a holiday — changes the result or who’s in it, rather than the treatment.
What randomisation already handles
Worth stating first, because the fear is usually larger than the risk.
Black Friday week: everyone's conversion rate doubles
week 1 (normal) week 2 (BF) pooled
control 3.0% 6.0% 4.5%
variant 3.3% 6.6% 4.95%
relative lift +10% +10% +10%
Both arms experience the season simultaneously, so the difference between them is unaffected. This is the same property that handles weather, news events and competitor activity — none of which you measured, all of which cancel — Why Randomisation Works.
So the standard worry — “our test ran over Christmas, is it invalid?” — usually has the answer no, it’s still internally valid.
What it does break
Three real problems, in descending order of how often they bite.
1. The result may not generalise. A test that ran entirely across Black Friday measured the effect on Black Friday shoppers, who are more motivated, more discount-driven and less price-sensitive to delivery charges than your usual traffic. Shipping that winner in February applies a peak-season finding to normal-season users.
test: removing the delivery-cost line from the basket summary
measured over Black Friday +6.2% shoppers already committed, deal-driven
same test, re-run in March +0.4% normal traffic notices delivery cost
Neither number is wrong. They’re answers to different questions, and only one of them is the question you asked — Post-Test Validation.
2. Effects genuinely differ by period. This is a real heterogeneous treatment effect with time as the dimension. Urgency messaging works during a sale and irritates in a quiet week; delivery-date estimates matter enormously in mid-December and barely in June.
3. Variance rises, so power falls. Peak weeks have larger day-to-day swings, and higher variance means wider intervals for the same sample. A test spanning a volatile period needs more traffic than the calculator assumed at a stable baseline — Statistical Power.
The one that does bias: changing the split mid-season
Seasonality plus a change in traffic allocation is a genuine bias, via Simpson’s Paradox.
week 1 (normal) 90/10 split control 90,000 @ 3.0% variant 10,000 @ 3.3%
week 2 (BF) 50/50 split control 100,000 @ 6.0% variant 100,000 @ 6.6%
pooled control 190,000 → 4.58%
variant 110,000 → 6.30%
pooled "lift" = +37% ← nonsense. the variant is simply weighted
more heavily towards the high-converting week
Each week shows +10%. Pooled shows +37%. The arms have different period mixes, so the pooled comparison is meaningless — Traffic Allocation.
Rules
- Full weeks, always. Weekday and weekend behaviour differ enough that a partial week reweights the sample. Nine days is worse than seven — Test Duration
- Don’t stop a test during an abnormal period, even if the rule fires. Let it run to the end of the anomaly so both arms carry the same mix
- Don’t start one the day a campaign launches unless the campaign is the context you want to measure in
- Never change the split mid-test. If you must, discard the pre-change period rather than pooling
- Record the calendar in the test’s record — campaigns, promotions, price changes, outages, competitor sales. A year later this is the only thing that explains an anomalous archived result — Annotation and Change Logs, Experiment Archive
- Re-test seasonal winners in a normal period before treating them as permanent, particularly anything about urgency, scarcity or discounting
Testing during peak at all
The commercial argument is usually “we can’t risk changes during our biggest week”, and the counter-argument is that peak is when a win is worth the most. Both are right; the resolution is about which tests, not whether.
DON'T test during peak TEST during peak
anything that could break changes to peak-specific
checkout — the downside is experiences: sale banners,
concentrated into days countdowns, delivery cut-off
messaging
anything you intend to ship
year-round, because the anything you'd only ever ship
result won't generalise during a peak anyway
low-powered tests hoping the guardrails watched much harder
extra traffic saves them than usual — the cost of harm
— it also raises variance per hour is at its maximum
The strongest argument for testing during peak is that peak-season experiences are otherwise shipped every year on nothing but opinion, and they’re the highest-value pages the business has.
Where it interacts
- Test Duration — the business-cycle constraint, of which weekly seasonality is the smallest instance
- Novelty and Primacy Effects — the other time-varying effect, and routinely confused with this one; novelty fades within a test, seasonality moves both arms together
- Holdout Groups — a long-running holdout measures cumulative effect across a full year of seasons, which no individual test can
- Seasonality — the analytics-side treatment of the same pattern, in reporting rather than testing