Alerting on Metrics
Date: 2026-08-17
Being told when a business number breaks, rather than finding out at the Monday meeting. It’s a different job from system alerting — the systems are all green, the pipeline is healthy, and conversion has been down 20% since Thursday because a tag stopped firing.
Metric alerting is an automated check on a business metric — orders, revenue, conversion, an event’s volume — that notifies someone when the value leaves its expected range.
Why system alerting doesn’t cover this
SYSTEM ALERTS SAY WHAT ACTUALLY HAPPENED
all services 200 OK a GTM change broke the purchase event
error rate normal → revenue reporting shows −100%
p99 latency fine → and nothing was "wrong"
uptime 100%
a payment method silently stopped
→ orders −18%, no errors logged,
the provider returns a clean decline
The infrastructure being healthy is not evidence that the business is. Metric alerting is the layer that watches outcomes rather than components — Alerting covers the system half, and the two need separate ownership because the people who can fix them are different.
What to alert on
Keep the list short. Ten to fifteen alerts for a commerce site, not fifty.
| Alert | Catches | Urgency |
|---|---|---|
| Orders per hour vs same hour last week | Broken checkout, payment outage, tracking break | Page |
| Any critical event volume at zero | A tag stopped firing | Page |
| Conversion rate by device | A bug affecting only mobile, invisible in the average | Page |
| Revenue per order out of range | Currency or units error — a 100× error is usually pence/pounds | Page |
| Traffic by channel vs baseline | Campaign stopped, tracking parameters dropped, an index drop | Ticket |
| Event volume by name, week on week | Taxonomy drift, a rename, a new event | Ticket |
| Tool-to-tool discrepancy widening | One pipeline degrading — Tool Discrepancies | Ticket |
| Data freshness / pipeline lag | The numbers are stale rather than wrong — Data Quality Monitoring | Page |
Zero-volume alerts are the highest-value ones on the list and the cheapest to build. A metric that flatlines needs no statistics, and a broken tag is the single most common cause of a business-metric emergency.
Threshold or deviation
STATIC THRESHOLD DEVIATION FROM EXPECTED
"orders/hr below 40" "orders/hr more than 3 robust SDs
below the same hour last week"
✓ trivial to build and explain ✓ survives day-of-week and season
✓ right for hard floors — zero, ✓ catches a 20% drop at 3am that a
or a known contractual minimum static floor never would
✗ fires every night and every ✗ needs history, and needs known
Sunday anomalies excluded from the baseline
✗ misses proportional drops in ✗ harder to explain to whoever gets
low-traffic hours woken up
Use static thresholds for zero and for absurd values; use deviation for everything else. The mechanics of the deviation test are in Anomaly Detection; the decisions here are which metrics and who responds.
Context is what makes it actionable
An alert saying “orders down 22%” produces a scramble. The same alert with context produces a diagnosis.
⚠ Orders/hour 22% below expected
observed 143 · expected 184 (same hour, prior 8 weeks) · 3.2 SD
↓ mobile −38% ← concentrated here
↓ desktop −2%
↓ Safari −61% ← and here
✓ add-to-cart normal → not traffic, not interest
✗ purchase events −40% → breaks between basket and confirmation
deploys in the last 4h: checkout-service v2.14 (2h ago)
campaigns changed: none
runbook: analytics-alerts/orders-drop
The two lines that resolve most incidents are the segment breakdown and the deploy list. Building them into the alert payload turns a 40-minute investigation into a 4-minute one — Annotation and Change Logs.
Alert fatigue, and the analytics-specific version
The general failure is in Alerting. What’s specific here:
- Business metrics move for business reasons. A campaign ending, a bank holiday, a competitor’s sale — all produce genuine deviations that need no action. This makes the false-positive rate structurally higher than for system alerts, and it’s why most of these should be tickets rather than pages
- Seasonality generates alerts on a schedule unless the baseline handles it — Seasonality
- Many metrics × frequent checks = daily false alarms from chance alone. Require persistence across two or three intervals before firing
- Nobody owns business metrics at 3am. An analytics alert routed to an on-call engineer who can’t interpret it gets acknowledged and forgotten. Route to someone who can act, and accept business-hours response for most of them
Measure the actionability rate. Of last month’s alerts, how many led to a change? Below half and people have already started ignoring the channel.
Making it survive
- Every alert has an owner and a runbook — what it means, how to confirm, what to check first
- Alert on absence, not just on movement. Zero is the highest-signal state and the easiest to miss
- Suppress during known events. Feed the campaign calendar and the deploy log into the alerting layer so planned changes don’t fire
- Review quarterly and delete. An alert that has fired eleven times and been actioned once is training people to ignore it
- Test it deliberately. An alert nobody has seen fire is a hypothesis — break something in a non-production environment and confirm it reaches a person
Where it interacts
- Anomaly Detection — the statistical method underneath the deviation thresholds
- Data Quality Monitoring — pipeline-level alerting, which catches problems before they reach a business metric
- Alerting — the system-side counterpart, with different owners and different urgency
- The symptom list — the ordered diagnosis once an alert has fired and is genuine