Tags: web-dev concept

Alerting

Date: 2026-08-17


Deciding what is worth waking someone for. The design constraint is human rather than technical: attention is finite and degrades with use, so every alert that fires without needing action makes the next real one less likely to be acted on.


Alerting is the set of rules that notify a person, by page or message, when a system’s measurements indicate something needs human action.

Symptoms, not causes

The rule that removes most bad alerts.

CAUSE ALERTS                         SYMPTOM ALERTS

CPU above 80%                        checkout success rate below 98%
disk 85% full                        p99 latency above 2s
memory climbing                      error rate above 1%
a pod restarted                      no orders in 10 minutes
replica lag 4s

fire constantly, often fine          fire when customers are affected
miss failures with normal CPU        catch causes you never predicted
one per resource, forever            a handful, stable

High CPU with a healthy site is not an incident. A healthy CPU with a broken checkout is. Cause metrics belong on dashboards, where they’re used for diagnosis after a symptom alert fires — Observability.

The narrow exception is predictable exhaustion: a disk filling at a rate that gives you four hours’ notice is worth a low-urgency ticket, because the symptom version of that alert arrives as an outage.

The four questions before creating one

  1. Is a customer affected, or about to be? If no — dashboard, not alert
  2. Is there an action? An alert with no runbook is a notification. If the response is “watch it”, it isn’t an alert
  3. Does it need to be now? Page versus ticket versus daily digest. Most things are not now
  4. Will it fire when nothing is wrong? Model the false positive rate honestly before shipping it

Thresholds that don’t cry wolf

Static thresholds on noisy metrics are the main source of false alarms. Three fixes, in increasing order of effort:

  • Duration. “Error rate above 1% for 5 minutes” rather than instantaneously. Removes almost all spike noise for the cost of five minutes’ detection delay
  • Burn rate against an error budget. Rather than a fixed threshold, alert on how fast the month’s allowance is being spent — a fast burn pages, a slow burn tickets. This is the approach worth adopting once SLOs exist — Service Level Objectives
  • Seasonality-aware baselines. Traffic at 3am Sunday is not traffic at 8pm Tuesday, and a fixed order-count threshold fires every night. Compare against the same hour last week — Seasonality

Alert on absence too. Zero errors, zero orders and zero events all look like health and are frequently the most severe failure available — a broken tag, a stopped consumer, a dead cron job — Webhooks, Instrumentation Debugging.

Alert fatigue

The failure mode, and it’s a ratchet: noisy alerts get muted, muted alerts stay muted, and the channel dies.

symptom                              what it means

"that one always fires"              delete it or fix it. today
alerts routed to a channel nobody
  has open                           it isn't an alert
> ~2 pages per on-call night         unsustainable; people leave
acknowledged without investigation   the alert has already stopped working
a filter rule in someone's inbox     the ratchet, complete

Measure the actionability rate — of alerts fired last month, how many led to a change. Below roughly half and the system is training people to ignore it. Deleting alerts is the most common correct fix and the one that feels most irresponsible.

Routing and escalation

  • Page for things needing action within minutes. Everything else is a ticket
  • Severity means response time, not how upset anyone is. Write it down: SEV1 = now, SEV2 = business hours, SEV3 = backlog
  • One owner per alert, named, and a rotation rather than an individual
  • Escalate automatically if unacknowledged. Relying on someone noticing is how alerts get missed at 4am
  • Group related alerts. One failing database that pages eleven times has buried its own cause
  • Every alert links to a runbook — what it means, how to confirm, first three things to try, who to escalate to. Written when the alert is created, not during the incident — Incident Response

Alerting on business metrics

Often the earliest and most reliable signal, because it’s the only one measuring the thing you actually care about.

  • Orders per minute against the same hour last week catches broken checkouts that throw no errors — a disabled payment method, a validation rule rejecting valid postcodes, a sold-out feed
  • Conversion rate by device or browser catches the bugs that only affect one segment and never move the average
  • Revenue anomalies catch currency and pricing faults that look fine in every technical metric — Revenue Metrics

The caveat: these alerts also fire for real business events — a campaign ending, a bank holiday, a competitor’s sale. They need context before anyone acts, which usually makes them tickets rather than pages — Annotation and Change Logs.

Where it interacts

  • Service Level Objectives — the principled basis for what’s worth alerting on, and the source of burn-rate alerts
  • Error Tracking — alert on new issues and rate spikes, never on individual errors
  • Incident Response — the alert is the start of the process, and the runbook is the bridge between them
  • Anomaly Detection — the statistical version, and the place where false positive rates need thinking about properly