Tags: statistics concept

Outliers and Robust Statistics

Date: 2026-08-17


Spotting the values distorting a result, and using estimators that don’t bend to them. The trap specific to detection is circular: the standard rule for finding outliers uses the mean and standard deviation, both of which the outliers have already corrupted.


An outlier is a value far enough from the rest to distort a summary of them; a robust statistic is one that barely moves when a few such values are added — the median rather than the mean.

The circularity

order values (£):  18, 22, 24, 25, 27, 29, 31, 34, 38, 4,200

mean  =  444.8          ← larger than 9 of the 10 values
sd    =  1,318

"mean ± 3sd"  =  −3,509  to  4,399
              →  4,200 is INSIDE the bounds
              →  the rule detects nothing

The single outlier inflated the standard deviation so much that it made itself look normal. This is why the mean-and-sd rule fails exactly when it’s needed.

The robust alternative uses statistics the outlier can’t move:

median  =  28                        (average of 27 and 29)

absolute deviations from median:
  10, 6, 4, 3, 1, 1, 3, 6, 10, 4,172

MAD  =  median of those  =  5        ← the 4,172 is just "the largest",
                                        it doesn't drag the median

scaled MAD  =  5 × 1.4826  =  7.41   ← makes it comparable to an sd
                                        for normal data

robust bounds, median ± 3 × scaled MAD
            =  28 ± 22.2   =   5.8  to  50.2
            →  4,200 is far outside. DETECTED

In plain terms: the median and MAD (median absolute deviation) describe the bulk of the data without being pulled by the extremes, so they can be used to identify the extremes. The mean and standard deviation cannot.

The estimators

StatisticBreakdown pointNotes
Mean0% — one value can move it anywhereWhat everyone reports
Median50%Half the data must be corrupted to break it
Trimmed mean (10%)10%Drop the top and bottom 10%, average the rest
Winsorised mean (10%)10%Replace them with the 10th/90th percentile, then average
Standard deviation0%Squares distances, so extremes dominate
IQR25%The middle 50%
MAD50%The robust spread measure

Breakdown point is the proportion of the data that has to be corrupted before the estimator becomes arbitrary. Zero for the mean means a single value is enough.

Worked on the same data:

mean                  £444.80    ← useless
median                 £28.00
10% trimmed mean       £28.75    (drop 18 and 4,200, average the rest)
10% winsorised mean    £30.50    (4,200 → 38, 18 → 22, then average)

Trim, winsorise, or keep

The decision matters and depends on what the outlier is.

IS IT AN ERROR?                      IS IT REAL BUT EXTREME?

£4,200 order that's actually         £4,200 order from a genuine
£42.00 with a currency bug           trade customer
  → FIX IT, or delete                  → it is part of your business
  → Revenue Metrics                → deleting it misstates revenue
                                       → but letting it decide a test
a test order, a bot, an                 is also wrong
internal transaction
  → exclude, on a rule fixed          → WINSORISE for analysis,
    in advance                          keep the real figure for finance

A currency-subunit bug is the commonest source of an apparent 100× outlier — Revenue Metrics.

Winsorisation is usually right for testing and wrong for reporting. A £5,000 trade order is real revenue and belongs in the P&L; letting it land in one arm of an A/B test and decide the outcome is measurement noise deciding a business question — Winsorisation and Capping.

The cap must be chosen before looking at the results. Choosing it afterwards, when you can see which arm the big order fell in, is P-Hacking — and it’s undetectable in a write-up. Fix the percentile at design time, apply it identically to both arms, and record it.

Detection methods

IQR RULE (Tukey)     outside  Q1 − 1.5×IQR  to  Q3 + 1.5×IQR
                     robust, standard, good default

MODIFIED Z-SCORE     |x − median| / (1.4826 × MAD)  >  3.5
                     robust, works on smaller samples

PERCENTILE CAP       anything above the 99th percentile
                     simple, no distributional assumption,
                     the usual practical choice for revenue

MEAN ± 3 SD          ✗ circular. don't

For heavy-tailed data, “outlier” is often the wrong frame. Revenue per visitor is genuinely heavy-tailed — the large values aren’t errors or anomalies, they’re the shape of the distribution. Nothing is contaminated; the mean is just a poor summary of it. In that case the answer is a different summary statistic or a transformation, not a detection rule — Skewed and Heavy-Tailed Distributions.

Where it changes decisions here

  • A/B tests on revenue. One large order can move an arm’s mean enough to flip a result. Capping is the standard defence and it substantially reduces variance too — which raises power at the same time — Metric Sensitivity
  • Performance metrics. Never a mean load time. One 40-second session on a bad connection distorts it; report p75 and p95 instead — Percentiles in Performance, Core Web Vitals
  • Correlation and Linear Regression. Both are extremely sensitive; a single point can create or destroy an apparent relationship. Refit without extreme points as a routine check
  • Average order value reporting. Median AOV and mean AOV can differ substantially, and the mean is the one everyone quotes — Average Order Value
  • Anomaly Detection. Baselines built from means and standard deviations get corrupted by the very spikes they should be catching

Where it interacts