Tags: statistics concept

Effect Size

Date: 2026-08-16


How big, as opposed to how certain. Significance testing answers whether an effect exists; only effect size answers whether it’s worth anything — and the relative/absolute confusion around it misprices tests by an order of magnitude.


What it is

Effect size is the magnitude of a difference, in units you can act on. It answers how much, where a p-value answers how sure.

The two are independent. A large effect can be uncertain and a trivial one can be nailed down precisely, which is why reporting one without the other is always incomplete.

Absolute and relative

The distinction that causes the most expensive mistakes in CRO, because both are called “percent”.

control  3.00%
variant  3.09%

absolute difference   3.09 − 3.00        = 0.09 percentage points  (pp)
relative difference   0.09 ÷ 3.00        = 3.0%

Percentage points measure the gap between two percentages. Percent measures the gap as a proportion of the starting value. A 3% relative lift on a 3% baseline is 0.09pp — and if you enter 3% into a sample size calculator’s absolute field, you’ll be told you need a few hundred visitors instead of half a million.

StatementReading
”Conversion up 3 percentage points”3.00% → 6.00%. Doubled
”Conversion up 3 percent”3.00% → 3.09%. Marginal

Always state which. Relative is the right default for talking about lifts — it’s comparable across pages with different baselines — but sample size and interval maths run on absolute, so the conversion happens constantly and quietly.

Which measure

Metric typeEffect sizeNotes
Conversion rateRelative lift, or difference in ppReport both; they answer different questions
Revenue per visitorAbsolute £, and relative %Absolute is what finance wants
Continuous, comparing across metricsCohen’s d — the difference in means divided by pooled standard deviationStandardised, so unitless and comparable. Rarely used in CRO but it’s what “small/medium/large effect” refers to

Cohen’s d worked, briefly: two groups differing by £0.15 in revenue per visitor with a pooled standard deviation of £13.44 gives d = 0.15 ÷ 13.44 = 0.011. By the usual rule of thumb — 0.2 small, 0.5 medium, 0.8 large — that’s a vanishingly small standardised effect. It’s also worth £75,000 a year on 500,000 visitors. Standardised effect sizes are a poor guide to commercial value, which is why ecommerce mostly ignores them.

Why it’s the number that matters

Three things follow from reporting the effect rather than the verdict:

  • It’s the input to the decision. Build cost, maintenance and risk are all in pounds. Significance isn’t in pounds
  • It’s what a p-value hides. The same p-value can come from a huge effect on small traffic or a trivial one on large traffic — P-Values
  • It’s what makes a result reusable. An archive recording “won, p=0.02” teaches nothing. One recording “+4.1% relative, CI +0.8% to +7.5%, on category pages” informs the next estimate

The observed effect is biased upward

An important and unintuitive one. If you only ship results that crossed the significance line, then among the effects you ship, the ones that got a lucky sample are over-represented. So the observed effect systematically overstates the true effect — and worse the less powered the test was.

In plain terms: you selected the winners on a noisy measurement, so the measurement was flattering by the act of selecting it. Expect the shipped change to underdeliver against its test result. This is the Winner’s Curse, and the fix is more power rather than more scepticism after the fact.

Practical

  • State the effect and its interval together. A point estimate alone implies a precision you don’t have
  • Convert to money before presenting. Annual visitors × baseline × relative lift × margin per order
  • Label pp and % explicitly, every time, including in calculator inputs
  • Compare effects to your MDE, not to zero. An effect below your MDE that came back significant should be treated with suspicion, not celebration