Tags: statistics concept
Outliers and Robust Statistics
Date: 2026-08-17
Spotting the values distorting a result, and using estimators that don’t bend to them. The trap specific to detection is circular: the standard rule for finding outliers uses the mean and standard deviation, both of which the outliers have already corrupted.
An outlier is a value far enough from the rest to distort a summary of them; a robust statistic is one that barely moves when a few such values are added — the median rather than the mean.
The circularity
order values (£): 18, 22, 24, 25, 27, 29, 31, 34, 38, 4,200
mean = 444.8 ← larger than 9 of the 10 values
sd = 1,318
"mean ± 3sd" = −3,509 to 4,399
→ 4,200 is INSIDE the bounds
→ the rule detects nothing
The single outlier inflated the standard deviation so much that it made itself look normal. This is why the mean-and-sd rule fails exactly when it’s needed.
The robust alternative uses statistics the outlier can’t move:
median = 28 (average of 27 and 29)
absolute deviations from median:
10, 6, 4, 3, 1, 1, 3, 6, 10, 4,172
MAD = median of those = 5 ← the 4,172 is just "the largest",
it doesn't drag the median
scaled MAD = 5 × 1.4826 = 7.41 ← makes it comparable to an sd
for normal data
robust bounds, median ± 3 × scaled MAD
= 28 ± 22.2 = 5.8 to 50.2
→ 4,200 is far outside. DETECTED
In plain terms: the median and MAD (median absolute deviation) describe the bulk of the data without being pulled by the extremes, so they can be used to identify the extremes. The mean and standard deviation cannot.
The estimators
| Statistic | Breakdown point | Notes |
|---|---|---|
| Mean | 0% — one value can move it anywhere | What everyone reports |
| Median | 50% | Half the data must be corrupted to break it |
| Trimmed mean (10%) | 10% | Drop the top and bottom 10%, average the rest |
| Winsorised mean (10%) | 10% | Replace them with the 10th/90th percentile, then average |
| Standard deviation | 0% | Squares distances, so extremes dominate |
| IQR | 25% | The middle 50% |
| MAD | 50% | The robust spread measure |
Breakdown point is the proportion of the data that has to be corrupted before the estimator becomes arbitrary. Zero for the mean means a single value is enough.
Worked on the same data:
mean £444.80 ← useless
median £28.00
10% trimmed mean £28.75 (drop 18 and 4,200, average the rest)
10% winsorised mean £30.50 (4,200 → 38, 18 → 22, then average)
Trim, winsorise, or keep
The decision matters and depends on what the outlier is.
IS IT AN ERROR? IS IT REAL BUT EXTREME?
£4,200 order that's actually £4,200 order from a genuine
£42.00 with a currency bug trade customer
→ FIX IT, or delete → it is part of your business
→ Revenue Metrics → deleting it misstates revenue
→ but letting it decide a test
a test order, a bot, an is also wrong
internal transaction
→ exclude, on a rule fixed → WINSORISE for analysis,
in advance keep the real figure for finance
A currency-subunit bug is the commonest source of an apparent 100× outlier — Revenue Metrics.
Winsorisation is usually right for testing and wrong for reporting. A £5,000 trade order is real revenue and belongs in the P&L; letting it land in one arm of an A/B test and decide the outcome is measurement noise deciding a business question — Winsorisation and Capping.
The cap must be chosen before looking at the results. Choosing it afterwards, when you can see which arm the big order fell in, is P-Hacking — and it’s undetectable in a write-up. Fix the percentile at design time, apply it identically to both arms, and record it.
Detection methods
IQR RULE (Tukey) outside Q1 − 1.5×IQR to Q3 + 1.5×IQR
robust, standard, good default
MODIFIED Z-SCORE |x − median| / (1.4826 × MAD) > 3.5
robust, works on smaller samples
PERCENTILE CAP anything above the 99th percentile
simple, no distributional assumption,
the usual practical choice for revenue
MEAN ± 3 SD ✗ circular. don't
For heavy-tailed data, “outlier” is often the wrong frame. Revenue per visitor is genuinely heavy-tailed — the large values aren’t errors or anomalies, they’re the shape of the distribution. Nothing is contaminated; the mean is just a poor summary of it. In that case the answer is a different summary statistic or a transformation, not a detection rule — Skewed and Heavy-Tailed Distributions.
Where it changes decisions here
- A/B tests on revenue. One large order can move an arm’s mean enough to flip a result. Capping is the standard defence and it substantially reduces variance too — which raises power at the same time — Metric Sensitivity
- Performance metrics. Never a mean load time. One 40-second session on a bad connection distorts it; report p75 and p95 instead — Percentiles in Performance, Core Web Vitals
- Correlation and Linear Regression. Both are extremely sensitive; a single point can create or destroy an apparent relationship. Refit without extreme points as a routine check
- Average order value reporting. Median AOV and mean AOV can differ substantially, and the mean is the one everyone quotes — Average Order Value
- Anomaly Detection. Baselines built from means and standard deviations get corrupted by the very spikes they should be catching
Where it interacts
- Winsorisation and Capping — the applied version, with the rules for doing it defensibly in a test
- Mean Median and Mode — three answers to “typical”, and when each one lies
- Skewed and Heavy-Tailed Distributions — where extreme values are the distribution rather than contamination of it
- Percentiles and Quantiles — the robust way to describe a distribution’s shape