Tags: experimentation concept

Novelty and Primacy Effects

Date: 2026-08-16


Two opposite reactions to change itself rather than to what changed. Novelty inflates the early result and fades; primacy depresses it and recovers. Both affect returning users and neither touches first-time visitors — which is also how you detect them.


What they are

Novelty effect — regular users notice something is different, engage with it because it’s new, and the lift decays as the novelty wears off.

Primacy effect — regular users have habits built around the old version, are disrupted by the change, and perform worse until they relearn. The result improves over time.

NOVELTY                              PRIMACY

lift                                 lift
 +8% ●                                 0% ────────────●───── settles
     ●                                −4%      ●    ●
 +4%    ●                             −8%  ●
 +1%       ●───●───● settles
     wk1  wk2  wk3  wk4                    wk1 wk2 wk3 wk4

ship on week 1 → underdelivers        kill on week 1 → threw away a winner

Both are real behaviour, not measurement error. The mistake is generalising a transient reaction into a permanent effect.

Why they only affect returning users

Neither is a reaction to the design — both are reactions to the difference between the design and what someone already knew. A first-time visitor has nothing to compare against, so they experience the variant as simply the site.

That gives you the detection method, and it’s the reliable one:

Split the result by new versus returning visitors.

  • Effect present in both, similar size → probably real
  • Effect much larger in returning → novelty
  • Effect negative in returning, neutral or positive in new → primacy
  • Effect only in new visitors → probably real, and the returning population is diluting it

This is one of the few segmentations worth pre-registering on every test, precisely because it’s diagnostic rather than exploratory — see Segmentation (test results).

The second detection method

Plot the effect over time, cumulatively and week by week. A stable effect is flat; novelty decays; primacy climbs.

Two cautions:

  • Weekly slices are underpowered. A nine-week test cut into weeks has ninth-sized samples each, so the noise is large. Read the trend, not any single week
  • This is not a licence to peek. Looking at the time trend after the test ends is analysis. Looking during it, and acting, is Peeking

What to do about them

  • Run long enough for the transient to pass. Two to four weeks is usually sufficient for novelty on a frequently-visited site; longer where visit frequency is low, because “week two” for the site is still “visit one” for the user
  • Weight the decision towards new visitors where the change is permanent and the returning population will eventually be habituated to it anyway
  • Post-Test Validation — check the shipped change is still delivering a month later. This is the only way to be sure, and it’s the step almost nobody does
  • Distinguish from Regression to the Mean, which produces the same fading pattern from a completely different cause. If the test was triggered by a page performing badly, regression is the more likely explanation, and it affects new and returning users alike

Where they bite hardest

  • High-frequency sites — daily or weekly visitors have strong habits, so primacy is pronounced. Retail with a long purchase cycle sees far less of either
  • Navigation and information architecture changes — primacy is severe, because habits are literally spatial. A navigation test that loses in week one routinely wins by week four
  • Promotional and visual changes — novelty is severe. A new banner style earns attention for being new
  • Loyalty and subscription audiences — the most habituated population you have

The honest summary: for a site with mostly new visitors, neither effect is likely to be your problem. For one with a loyal returning base, both are, and the new/returning split should be standard on every readout.