Tags: statistics concept

Regression to the Mean

Date: 2026-08-16


Extreme measurements contain more than their share of luck, and luck doesn’t repeat. So anything selected for being extreme moves back towards average on its own — and whatever you did in between gets the credit.


What it is

Regression to the mean is the tendency for an extreme measurement to be followed by a less extreme one, purely because part of what made it extreme was luck.

Nothing has to happen for it to occur. It appears whenever you measure something twice and select on the first measurement.

The mechanism

Any observed value is signal plus noise:

observed = true value + noise

Selecting the most extreme observations selects for both — the genuinely extreme and the ones that got an unlucky roll. On the next measurement the true value persists and the noise is redrawn, so the group’s average moves back towards the middle.

Nothing causes this. It’s a property of measuring twice.

Worked

Twelve category pages, all with a true conversion rate of exactly 3.0%. Any variation you see is noise. You pick the worst five, redesign them, and measure again.

                 period 1   → selected? →   period 2
page A             2.1%          ✓            2.9%
page B             2.3%          ✓            3.2%
page C             2.4%          ✓            2.8%
page D             2.6%          ✓            3.4%
page E             2.7%          ✓            3.0%
page F–L      3.0% – 3.9%        ✗              —

worst five, mean period 1                    2.42%
worst five, mean period 2                    3.06%
                                             ─────
apparent improvement                         +26% relative

Every page had a true rate of 3.0% throughout, and the redesign did nothing. The +26% is entirely the selection.

In plain terms: you chose the pages that were having a bad month. Bad months end whether or not anyone intervenes.

Why CRO is unusually exposed

The standard workflow is find the worst-performing thing and fix it. That is precisely the selection step that manufactures the illusion — and the weaker the page’s traffic, the noisier its measurement, so the pages you’re most likely to select are the ones where the effect is largest.

The same shape recurs:

  • Optimise the worst-converting product pages → they improve
  • Intervene on the customers most likely to churn → churn falls in that group
  • Fix the slowest pages → they get faster
  • Rescue an underperforming campaign → performance recovers

In each case a real intervention may also have helped. The problem is that you cannot tell how much, because a null intervention would have produced an improvement too.

Distinguishing it from a real effect

SymptomMore likely
Improvement appears in the selected group onlyRegression
Improvement holds against a comparison group selected the same way but untreatedReal effect
Effect fades on regulars over weeks, was strongest in week oneNovelty and Primacy Effects
Shipped effect smaller than the test measuredWinner’s Curse — same mechanism, applied to test results rather than pages

The only reliable defence is a control group selected identically. Pick the worst ten, treat five at random, leave five. Both regress; the difference between them is the effect. That’s Randomised Controlled Trials applied to a selection problem, and it’s why an A/B test on a badly-performing page is trustworthy while a before/after on the same page is not.

Elsewhere

  • The winner’s curse — shipped test winners underdeliver, because you selected the ones with favourable noise
  • Sports and business punditry — the “second season slump” after an exceptional debut
  • Any before/after comparison triggered by a bad number. If the reason you’re measuring is that something looked bad, expect improvement regardless
  • Alerting — anomaly alerts fire on extremes, so the next reading usually looks like recovery. See Anomaly Detection

Related but distinct from Selection Bias: that’s about who ends up in the sample, this is about what happens when you measure the same units twice.