Tags: statistics concept

The Central Limit Theorem

Date: 2026-08-16


Sample averages become normally distributed even when the underlying data is nothing like normal. It’s the result that makes standard testing work on messy real data — and the conditions on it are exactly where revenue metrics fail.


What it is

The central limit theorem (CLT) states that the distribution of sample means approaches a normal distribution as sample size grows, regardless of the shape of the population the samples are drawn from.

POPULATION                      DISTRIBUTION OF SAMPLE MEANS

conversions: 0 or 1             take 1,000 users, compute the rate
▌                    ▌          repeat many times
▌                    ▌
──────────────────────                ╱▔▔╲
  not remotely normal                ╱    ╲
                                   ─╯      ╰─
                                approximately normal

Why it’s the load-bearing result

Every standard test computes something like “how far is my observed difference from zero, in standard errors, and how surprising is that?” — and answering the second half requires knowing the distribution of the statistic.

The CLT supplies it. You don’t need to know how order values or conversions are distributed; you need the average to be approximately normal, and the CLT says it will be.

In plain terms: individual customers are chaotic and their average is well-behaved. Statistics works on the average, which is why it works at all.

The conditions, and where they fail

Three, and the third is the one that matters commercially.

Independence. Observations must not influence each other. Violated by counting the same user twice — Randomisation Unit.

Finite variance. True of everything you’ll meet in practice.

Enough samples — and “enough” depends on skew. This is the practical constraint:

Underlying dataSamples for the mean to be near-normal
Symmetric~30
Mildly skewed (conversion rates)Hundreds
Heavily skewed (revenue per visitor)Thousands, sometimes tens of thousands

Revenue is where this bites. Order value is heavy-tailed, so the sample mean converges slowly. At modest sample sizes the normal approximation is poor and the resulting confidence intervals are too narrow — they understate uncertainty, which is the direction that produces false confidence rather than caution.

That’s the technical reason revenue per visitor is hard to test on, sitting underneath the practical one about variance in Metric Sensitivity.

What it does not say

Common misreadings, each of which causes real errors:

  • It says nothing about your raw data becoming normal. More customers doesn’t make order values bell-shaped. It only concerns the distribution of the mean
  • It doesn’t fix outliers. A heavy tail still drags the mean, and the CLT concerns the shape of the sampling distribution, not whether the mean is a sensible summary — Outliers and Robust Statistics
  • It doesn’t apply to percentiles. The CLT is about means. Confidence intervals for a median or a p75 need different machinery — usually Bootstrapping
  • It doesn’t rescue a biased sample. It describes sampling variation, not systematic error. A biased sample of a million is still biased — Selection Bias

What to do when it doesn’t hold

In order of preference:

  1. Bootstrapping — resample your own data to build the sampling distribution empirically. No distributional assumption at all, and it’s the honest answer for revenue
  2. Winsorisation and Capping — cap the tail so the mean converges faster. Pre-registered threshold, applied to both arms
  3. Change the metric. Conversion rate has no tail. Often the real answer is that revenue per visitor was the wrong primary — Overall Evaluation Criterion
  4. A non-parametric test, comparing distributions rather than means — Other Statistical Tests

The practical test: plot the metric. If a small number of observations account for a large share of the total, the tail is heavy enough to worry about, and a normal-approximation interval will be too confident.