Tags: analytics concept

Alerting on Metrics

Date: 2026-08-17


Being told when a business number breaks, rather than finding out at the Monday meeting. It’s a different job from system alerting — the systems are all green, the pipeline is healthy, and conversion has been down 20% since Thursday because a tag stopped firing.


Metric alerting is an automated check on a business metric — orders, revenue, conversion, an event’s volume — that notifies someone when the value leaves its expected range.

Why system alerting doesn’t cover this

SYSTEM ALERTS SAY              WHAT ACTUALLY HAPPENED

all services 200 OK            a GTM change broke the purchase event
error rate normal              → revenue reporting shows −100%
p99 latency fine               → and nothing was "wrong"
uptime 100%
                               a payment method silently stopped
                               → orders −18%, no errors logged,
                                 the provider returns a clean decline

The infrastructure being healthy is not evidence that the business is. Metric alerting is the layer that watches outcomes rather than components — Alerting covers the system half, and the two need separate ownership because the people who can fix them are different.

What to alert on

Keep the list short. Ten to fifteen alerts for a commerce site, not fifty.

AlertCatchesUrgency
Orders per hour vs same hour last weekBroken checkout, payment outage, tracking breakPage
Any critical event volume at zeroA tag stopped firingPage
Conversion rate by deviceA bug affecting only mobile, invisible in the averagePage
Revenue per order out of rangeCurrency or units error — a 100× error is usually pence/poundsPage
Traffic by channel vs baselineCampaign stopped, tracking parameters dropped, an index dropTicket
Event volume by name, week on weekTaxonomy drift, a rename, a new eventTicket
Tool-to-tool discrepancy wideningOne pipeline degrading — Tool DiscrepanciesTicket
Data freshness / pipeline lagThe numbers are stale rather than wrong — Data Quality MonitoringPage

Zero-volume alerts are the highest-value ones on the list and the cheapest to build. A metric that flatlines needs no statistics, and a broken tag is the single most common cause of a business-metric emergency.

Threshold or deviation

STATIC THRESHOLD                    DEVIATION FROM EXPECTED

"orders/hr below 40"                "orders/hr more than 3 robust SDs
                                     below the same hour last week"

✓ trivial to build and explain      ✓ survives day-of-week and season
✓ right for hard floors — zero,     ✓ catches a 20% drop at 3am that a
  or a known contractual minimum      static floor never would
✗ fires every night and every       ✗ needs history, and needs known
  Sunday                              anomalies excluded from the baseline
✗ misses proportional drops in      ✗ harder to explain to whoever gets
  low-traffic hours                    woken up

Use static thresholds for zero and for absurd values; use deviation for everything else. The mechanics of the deviation test are in Anomaly Detection; the decisions here are which metrics and who responds.

Context is what makes it actionable

An alert saying “orders down 22%” produces a scramble. The same alert with context produces a diagnosis.

⚠ Orders/hour 22% below expected
  observed 143 · expected 184 (same hour, prior 8 weeks) · 3.2 SD

  ↓ mobile     −38%    ← concentrated here
  ↓ desktop     −2%
  ↓ Safari     −61%    ← and here
  ✓ add-to-cart normal        → not traffic, not interest
  ✗ purchase events −40%      → breaks between basket and confirmation

  deploys in the last 4h:  checkout-service v2.14 (2h ago)
  campaigns changed:       none
  runbook: analytics-alerts/orders-drop

The two lines that resolve most incidents are the segment breakdown and the deploy list. Building them into the alert payload turns a 40-minute investigation into a 4-minute one — Annotation and Change Logs.

Alert fatigue, and the analytics-specific version

The general failure is in Alerting. What’s specific here:

  • Business metrics move for business reasons. A campaign ending, a bank holiday, a competitor’s sale — all produce genuine deviations that need no action. This makes the false-positive rate structurally higher than for system alerts, and it’s why most of these should be tickets rather than pages
  • Seasonality generates alerts on a schedule unless the baseline handles it — Seasonality
  • Many metrics × frequent checks = daily false alarms from chance alone. Require persistence across two or three intervals before firing
  • Nobody owns business metrics at 3am. An analytics alert routed to an on-call engineer who can’t interpret it gets acknowledged and forgotten. Route to someone who can act, and accept business-hours response for most of them

Measure the actionability rate. Of last month’s alerts, how many led to a change? Below half and people have already started ignoring the channel.

Making it survive

  • Every alert has an owner and a runbook — what it means, how to confirm, what to check first
  • Alert on absence, not just on movement. Zero is the highest-signal state and the easiest to miss
  • Suppress during known events. Feed the campaign calendar and the deploy log into the alerting layer so planned changes don’t fire
  • Review quarterly and delete. An alert that has fired eleven times and been actioned once is training people to ignore it
  • Test it deliberately. An alert nobody has seen fire is a hypothesis — break something in a non-production environment and confirm it reaches a person

Where it interacts

  • Anomaly Detection — the statistical method underneath the deviation thresholds
  • Data Quality Monitoring — pipeline-level alerting, which catches problems before they reach a business metric
  • Alerting — the system-side counterpart, with different owners and different urgency
  • The symptom list — the ordered diagnosis once an alert has fired and is genuine