Tags: statistics concept

Bootstrapping

Date: 2026-08-16


Build the confidence interval by resampling your own data instead of assuming a distribution. It’s the honest answer for revenue and for percentiles, and it works by simulating the repeated experiments you can’t actually run.


What it is

Bootstrapping estimates the sampling distribution of a statistic by repeatedly resampling with replacement from the data you have, recomputing the statistic each time.

observed data      n = 53,000 revenue-per-visitor values
                          │
       ┌──────────────────┼──────────────────┐
       ▼                  ▼                  ▼
  resample n         resample n         resample n     × 10,000
  with replacement   with replacement   with replacement
       │                  │                  │
   mean £1.48         mean £1.53         mean £1.51
                          │
              10,000 means, forming a distribution
                          │
        take the 2.5th and 97.5th percentiles
              → 95% confidence interval

With replacement is the crucial detail. Each resample is the same size as the original and may include some observations twice and others not at all — that variation is what simulates drawing a fresh sample from the population.

Why it works

Standard confidence intervals assume you know the shape of the sampling distribution — usually normal, via The Central Limit Theorem. Bootstrapping observes it instead of assuming it.

In plain terms: you can’t rerun the experiment a thousand times, so you rerun it a thousand times against your own data. If a few enormous orders are driving the result, some resamples will include them and some won’t, and the resulting spread tells you honestly how unstable the estimate is.

When it’s the right tool

  • Heavy-tailed metrics. Revenue per visitor, order value — where the normal approximation converges slowly and produces intervals that are too narrow — Skewed and Heavy-Tailed Distributions
  • Percentiles. The CLT is about means; there’s no simple formula for the standard error of a median or a p75. Bootstrapping is the standard route — Percentiles and Quantiles
  • Ratio metrics, where the variance of a ratio isn’t the variance of its parts — Ratio Metrics
  • Any statistic with no closed-form standard error — a difference in medians, a Gini coefficient, anything custom

What it doesn’t fix

Worth being clear, because it gets over-claimed:

  • It doesn’t fix bias. Resampling a biased sample gives you a precise interval around a wrong number — Selection Bias
  • It doesn’t fix small samples. With 30 observations you’re resampling 30 observations; the bootstrap can’t manufacture information that isn’t there
  • It doesn’t fix dependence. If observations aren’t independent — the same user twice — the resample inherits that
  • The tail is still under-represented. If the true distribution has a 1-in-10,000 event and your sample doesn’t contain one, no resampling will invent it

Practically

  • 10,000 resamples is a conventional default; a few thousand is usually enough. It’s cheap — this is a compute problem, not a statistics problem
  • Bootstrap the difference, not each arm separately. You want an interval on the effect
  • Percentile method — take the 2.5th and 97.5th percentiles of the bootstrap distribution — is the simple version and adequate for most work. Bias-corrected variants exist and matter at the margins
  • Set a seed so the result is reproducible. An interval that changes each time you run it is hard to defend

Where it fits

Most experimentation platforms use a normal approximation because it’s fast and closed-form, and it’s fine for conversion rates. For revenue metrics it’s frequently optimistic about precision.

If your platform reports a revenue interval and you can get the raw data, bootstrapping it yourself is a worthwhile check — a materially wider bootstrap interval means the reported one was understating uncertainty, and that changes what the result supports. See Guide - Statistics for CRO.