Tags: experimentation concept
Novelty and Primacy Effects
Date: 2026-08-16
Two opposite reactions to change itself rather than to what changed. Novelty inflates the early result and fades; primacy depresses it and recovers. Both affect returning users and neither touches first-time visitors — which is also how you detect them.
What they are
Novelty effect — regular users notice something is different, engage with it because it’s new, and the lift decays as the novelty wears off.
Primacy effect — regular users have habits built around the old version, are disrupted by the change, and perform worse until they relearn. The result improves over time.
NOVELTY PRIMACY
lift lift
+8% ● 0% ────────────●───── settles
● −4% ● ●
+4% ● −8% ●
+1% ●───●───● settles
wk1 wk2 wk3 wk4 wk1 wk2 wk3 wk4
ship on week 1 → underdelivers kill on week 1 → threw away a winner
Both are real behaviour, not measurement error. The mistake is generalising a transient reaction into a permanent effect.
Why they only affect returning users
Neither is a reaction to the design — both are reactions to the difference between the design and what someone already knew. A first-time visitor has nothing to compare against, so they experience the variant as simply the site.
That gives you the detection method, and it’s the reliable one:
Split the result by new versus returning visitors.
- Effect present in both, similar size → probably real
- Effect much larger in returning → novelty
- Effect negative in returning, neutral or positive in new → primacy
- Effect only in new visitors → probably real, and the returning population is diluting it
This is one of the few segmentations worth pre-registering on every test, precisely because it’s diagnostic rather than exploratory — see Segmentation (test results).
The second detection method
Plot the effect over time, cumulatively and week by week. A stable effect is flat; novelty decays; primacy climbs.
Two cautions:
- Weekly slices are underpowered. A nine-week test cut into weeks has ninth-sized samples each, so the noise is large. Read the trend, not any single week
- This is not a licence to peek. Looking at the time trend after the test ends is analysis. Looking during it, and acting, is Peeking
What to do about them
- Run long enough for the transient to pass. Two to four weeks is usually sufficient for novelty on a frequently-visited site; longer where visit frequency is low, because “week two” for the site is still “visit one” for the user
- Weight the decision towards new visitors where the change is permanent and the returning population will eventually be habituated to it anyway
- Post-Test Validation — check the shipped change is still delivering a month later. This is the only way to be sure, and it’s the step almost nobody does
- Distinguish from Regression to the Mean, which produces the same fading pattern from a completely different cause. If the test was triggered by a page performing badly, regression is the more likely explanation, and it affects new and returning users alike
Where they bite hardest
- High-frequency sites — daily or weekly visitors have strong habits, so primacy is pronounced. Retail with a long purchase cycle sees far less of either
- Navigation and information architecture changes — primacy is severe, because habits are literally spatial. A navigation test that loses in week one routinely wins by week four
- Promotional and visual changes — novelty is severe. A new banner style earns attention for being new
- Loyalty and subscription audiences — the most habituated population you have
The honest summary: for a site with mostly new visitors, neither effect is likely to be your problem. For one with a loyal returning base, both are, and the new/returning split should be standard on every readout.