Tags: statistics experimentation concept
Winsorisation and Capping
Date: 2026-08-16
Replace extreme values with a ceiling so one £5,000 order can’t decide a test. It’s a legitimate variance-reduction technique and a laundering mechanism, and the only thing separating them is whether you set the threshold before looking.
Winsorise or trim
Two different operations, routinely confused.
| What it does to a value above the cap | Sample size | Bias | |
|---|---|---|---|
| Winsorising | Replaces it with the cap | Unchanged | Pulls the mean down; keeps the observation |
| Trimming | Deletes the row entirely | Shrinks | Pulls the mean down; discards information |
raw order values, 99th percentile = £400
£42 £58 £31 £127 £5,000 £73
winsorised at £400: £42 £58 £31 £127 £400 £73 ← 6 values
trimmed at £400: £42 £58 £31 £127 £73 ← 5 values
Winsorising is almost always the right one in experimentation. Trimming removes a customer who genuinely bought, which changes the population you’re describing and breaks the denominator on any rate calculated alongside.
Why it’s worth doing
Revenue per visitor’s variance is dominated by order-value spread, and sample size scales with variance. From Metric Sensitivity’s example — 3% conversion, £50 mean order, 10% relative minimum detectable effect:
σ of order value = £60 σ of RPV = £13.44 → 126,000 per arm
σ of order value = £40 σ of RPV = £10.99 → 84,000 per arm
Working for the second row:
E[X²] = 0.03 × (50² + 40²) = 0.03 × 4,100 = 123
Var = 123 − 1.50² = 120.75
σ = √120.75 = £10.99
n = 2 × 7.849 × 120.75 ÷ (0.15)² = 84,200
A third less traffic, for capping the top 1% of orders. On a nine-week test that’s three weeks back.
Two caveats on that figure. Capping also pulls the mean down slightly, which shrinks the absolute effect you’re detecting and claws back some of the saving — so treat it as indicative. And the saving depends entirely on how heavy your tail is; a site with tight order values has nothing to gain.
The rules that keep it honest
In plain terms: capping decides which customers count fully. Decide that before you know who they are, or you’re choosing an answer rather than a method.
- Set the threshold before looking at results. Written into the analysis plan alongside the metric — Pre-Registration
- Derive it from historical data, not from this test’s data. Last quarter’s 99th percentile, computed once
- Apply it identically to both arms. Capping the variant’s outliers and not control’s is the most direct form of fraud available in experimentation
- Report both. Capped as the primary result, uncapped alongside. If they disagree materially, that disagreement is the finding
- Choose a percentile, not a value. £400 becomes wrong as the business changes; the 99th percentile stays meaningful
When it hides something real
Capping assumes the tail is noise. Sometimes it’s the effect.
- A change targeting high-value customers — a bulk-order flow, a trade pricing tier, a loyalty tier. If the variant works by making large orders larger, capping deletes exactly the effect you were testing
- The variant changed the tail’s shape. If control has three £5,000 orders and the variant has fifteen, that’s a result, not an outlier problem
- Business-to-business or trade accounts in an otherwise consumer basket. Two populations sharing a metric — segment them rather than capping across both
The check: look at the distribution above the cap in each arm before applying it. If the counts differ substantially, capping is concealing rather than de-noising.
Alternatives
- Variance Reduction — CUPED (controlled experiment using pre-experiment data) uses pre-period behaviour and doesn’t touch the data. Prefer it where users have history
- Bootstrapping — no distributional assumption, so no need to force the data to behave. Slower, and honest about heavy tails
- A different metric — conversion rate has no tail at all. Often the real answer is that revenue per visitor was the wrong primary
- Log transformation — compresses the tail without discarding it, at the cost of a result reported in units nobody can interpret commercially