Tags: experimentation concept

Seasonality in Tests

Date: 2026-08-17


Running across a peak, a payday or a campaign, and mistaking the period for the treatment. Randomisation protects the comparison — both arms get the same Black Friday — so seasonality rarely biases a result. What it does is change who the result applies to, which is a generalisation problem rather than a validity one.


Seasonality in tests is the risk that the period a test runs in — a sale, payday, a holiday — changes the result or who’s in it, rather than the treatment.

What randomisation already handles

Worth stating first, because the fear is usually larger than the risk.

Black Friday week: everyone's conversion rate doubles

              week 1 (normal)    week 2 (BF)     pooled
control            3.0%              6.0%          4.5%
variant            3.3%              6.6%          4.95%

relative lift     +10%              +10%          +10%

Both arms experience the season simultaneously, so the difference between them is unaffected. This is the same property that handles weather, news events and competitor activity — none of which you measured, all of which cancel — Why Randomisation Works.

So the standard worry — “our test ran over Christmas, is it invalid?” — usually has the answer no, it’s still internally valid.

What it does break

Three real problems, in descending order of how often they bite.

1. The result may not generalise. A test that ran entirely across Black Friday measured the effect on Black Friday shoppers, who are more motivated, more discount-driven and less price-sensitive to delivery charges than your usual traffic. Shipping that winner in February applies a peak-season finding to normal-season users.

test: removing the delivery-cost line from the basket summary

measured over Black Friday   +6.2%    shoppers already committed, deal-driven
same test, re-run in March   +0.4%    normal traffic notices delivery cost

Neither number is wrong. They’re answers to different questions, and only one of them is the question you asked — Post-Test Validation.

2. Effects genuinely differ by period. This is a real heterogeneous treatment effect with time as the dimension. Urgency messaging works during a sale and irritates in a quiet week; delivery-date estimates matter enormously in mid-December and barely in June.

3. Variance rises, so power falls. Peak weeks have larger day-to-day swings, and higher variance means wider intervals for the same sample. A test spanning a volatile period needs more traffic than the calculator assumed at a stable baseline — Statistical Power.

The one that does bias: changing the split mid-season

Seasonality plus a change in traffic allocation is a genuine bias, via Simpson’s Paradox.

week 1 (normal)   90/10 split   control 90,000 @ 3.0%   variant 10,000 @ 3.3%
week 2 (BF)       50/50 split   control 100,000 @ 6.0%  variant 100,000 @ 6.6%

pooled            control 190,000 → 4.58%
                  variant 110,000 → 6.30%

pooled "lift" = +37%     ← nonsense. the variant is simply weighted
                            more heavily towards the high-converting week

Each week shows +10%. Pooled shows +37%. The arms have different period mixes, so the pooled comparison is meaningless — Traffic Allocation.

Rules

  • Full weeks, always. Weekday and weekend behaviour differ enough that a partial week reweights the sample. Nine days is worse than seven — Test Duration
  • Don’t stop a test during an abnormal period, even if the rule fires. Let it run to the end of the anomaly so both arms carry the same mix
  • Don’t start one the day a campaign launches unless the campaign is the context you want to measure in
  • Never change the split mid-test. If you must, discard the pre-change period rather than pooling
  • Record the calendar in the test’s record — campaigns, promotions, price changes, outages, competitor sales. A year later this is the only thing that explains an anomalous archived result — Annotation and Change Logs, Experiment Archive
  • Re-test seasonal winners in a normal period before treating them as permanent, particularly anything about urgency, scarcity or discounting

Testing during peak at all

The commercial argument is usually “we can’t risk changes during our biggest week”, and the counter-argument is that peak is when a win is worth the most. Both are right; the resolution is about which tests, not whether.

DON'T test during peak                TEST during peak

anything that could break             changes to peak-specific
  checkout — the downside is           experiences: sale banners,
  concentrated into days               countdowns, delivery cut-off
                                       messaging
anything you intend to ship
  year-round, because the             anything you'd only ever ship
  result won't generalise              during a peak anyway

low-powered tests hoping the          guardrails watched much harder
  extra traffic saves them             than usual — the cost of harm
  — it also raises variance            per hour is at its maximum

The strongest argument for testing during peak is that peak-season experiences are otherwise shipped every year on nothing but opinion, and they’re the highest-value pages the business has.

Where it interacts

  • Test Duration — the business-cycle constraint, of which weekly seasonality is the smallest instance
  • Novelty and Primacy Effects — the other time-varying effect, and routinely confused with this one; novelty fades within a test, seasonality moves both arms together
  • Holdout Groups — a long-running holdout measures cumulative effect across a full year of seasons, which no individual test can
  • Seasonality — the analytics-side treatment of the same pattern, in reporting rather than testing