Tags: analytics statistics concept

Anomaly Detection

Date: 2026-08-17


Separating a genuine break from ordinary variation. The hard part isn’t the statistics — it’s that the two things that make commerce data anomalous, seasonality and heavy tails, both break the standard method, so a naive z-score alerts every Sunday and misses the actual outage.


Anomaly detection is flagging a data point that falls outside what the series’ own history predicts for that moment — not outside a fixed threshold.

The naive approach, and why it fails

orders per hour, threshold at mean ± 3 standard deviations

mean over 30 days   184
sd                   96
bounds           −104 to 472

  → the lower bound is negative. it can never fire.
  → the upper bound fires every Sunday evening AND every Black Friday
  → an outage at 3am, when normal is 12/hr, is invisible inside a
    band built from daytime volume

Two separate failures. The distribution isn’t normal — order counts are skewed and bounded below at zero. And the mean isn’t stationary — it moves by hour, by weekday, by season, so a single mean describes nothing.

Removing the pattern first

The whole method is: model what’s expected, then test the residual.

1  DECOMPOSE      strip out trend and seasonality — Seasonality
2  RESIDUAL       observed − expected
3  TEST           is this residual unusual for THIS series?
4  CONTEXTUALISE  is there a known reason? — Annotation and Change Logs
// same hour, same weekday, previous 8 weeks — the cheapest thing that works
function isAnomalous(series, now, k = 3) {
  const comparable = series.filter(p =>
    p.hour === now.hour && p.weekday === now.weekday
  ).slice(-8);
 
  const values = comparable.map(p => p.value).sort((a, b) => a - b);
  const median = values[Math.floor(values.length / 2)];
 
  // median absolute deviation — robust, unlike sd, to the outliers
  // that are exactly what you're trying to detect
  const mad = values.map(v => Math.abs(v - median)).sort((a, b) => a - b)[
    Math.floor(values.length / 2)
  ];
  const scale = mad * 1.4826;          // makes MAD comparable to a normal sd
 
  return { anomalous: Math.abs(now.value - median) > k * scale, median, scale };
}

Step 4 is the one that saves the most time: before investigating, check whether a deploy, campaign or outage already explains it — Annotation and Change Logs.

Why median and MAD rather than mean and standard deviation: a single past spike inflates the standard deviation, which widens the band, which stops future spikes being detected. The metric you use to find outliers must not itself be moved by outliers — Outliers and Robust Statistics, Mean Median and Mode.

Worked. Same-hour-same-weekday values over 8 weeks: 168, 172, 175, 179, 181, 184, 190, 240.

median                     180
absolute deviations        12, 8, 5, 1, 1, 4, 10, 60
MAD (median of those)      6
scale = 6 × 1.4826         8.9
bounds at k=3              180 ± 26.7  →  153 to 207

today's value 148  →  anomalous
                      note the 240 spike didn't widen the band;
                      with sd it would have (sd ≈ 23, bounds 111–249,
                      and 148 would have passed unnoticed)

The methods, ranked by what they cost

MethodHandlesCost
Same period last weekWeekly seasonalityTrivial. Start here, it solves most cases
Robust z-score on residuals (above)Weekly + outlier contaminationLow. The sensible default
STL decomposition — seasonal-trend using LoessMultiple seasonalities, changing trendModerate; needs a library and history
Prophet / ARIMA-family forecastingHolidays, multiple cycles, growthHigher; needs tuning and a person who understands it
Learned / ML detectionComplex multivariate patternsHigh, opaque, and hard to debug at 3am

Almost nobody needs the bottom two rows. The value curve is steep at the top and flat after — most missed anomalies are missed because nothing was watching the metric, not because the model was insufficiently sophisticated.

The two errors, priced

                     actually broken       actually fine
alert fires          ✓ caught              ✗ false positive
                                             → costs attention
no alert             ✗ MISSED               ✓ quiet
                       → costs revenue,
                         often for days

Set k from the cost ratio, not from convention. For tracking breaks — where a missed anomaly means days of unusable data and an unrecoverable gap — a lower threshold with more false positives is correct. For a metric where the response is “someone investigates for an hour”, the reverse.

And with many metrics watched, multiplicity bites hard:

40 metrics × 24 hourly checks = 960 tests/day
at a 1% false positive rate    ≈ 10 false alarms per day

In plain terms: watch enough metrics often enough and you will get alerts every day from pure chance. This is the arithmetic behind alert fatigue, and it’s why “alert on everything” collapses within a fortnight — The Multiple Comparisons Problem, Alerting on Metrics.

Mitigations that work: require persistence (anomalous for 2–3 consecutive intervals), require magnitude as well as significance, and group related metrics so one root cause produces one alert.

The anomalies worth watching for

  • Zero, or near-zero. A metric that flatlines is the highest-signal anomaly available and needs no statistics — a tag stopped firing, a consumer died, a job didn’t run
  • Sudden step changes, which almost always mean a deploy or a config change rather than customer behaviour — Metric Drift
  • Cardinality explosions — a new event name, a URL parameter appearing in a dimension — Cardinality
  • Ratio breaks with stable components. Sessions flat, users flat, sessions-per-user moved: that’s a definitional or identity change — Sessionisation
  • Distribution shifts with a stable mean. The average holds while the shape changes underneath, which no threshold on the mean will catch — Percentiles and Quantiles

Where it interacts

  • Seasonality — the pattern that must be removed before any of this means anything
  • Alerting on Metrics — what you do with a detection, and the human side of thresholds
  • Data Quality Monitoring — anomaly detection applied to the pipeline rather than to the business number
  • The symptom list — once something is genuinely anomalous, the ordered list of what to check