Observability
Date: 2026-08-17
Being able to answer questions about a running system that nobody thought to ask in advance. Monitoring tells you the thing you predicted has happened; observability lets you investigate the thing you didn’t predict — which is every interesting incident.
Observability is the degree to which a system’s internal state can be inferred from its external outputs — logs, metrics and traces — including for questions not anticipated when the instrumentation was written.
Monitoring versus observability
MONITORING OBSERVABILITY
"is CPU above 80%" "why are Belgian customers on Safari
"is the error rate above 1%" seeing checkout fail, but only when
"is the site up" they have a gift card applied"
known questions, asked repeatedly unknown questions, asked once
dashboards you built in advance queries you write during the incident
predicted failures the ones you didn't predict
You need both. Monitoring is what wakes you up; observability is what you use once you’re awake.
The three signals
Each answers a different question, and using one for another’s job is the common mistake.
| Metrics | Logs | Traces | |
|---|---|---|---|
| Question | Is something wrong, and how much? | What exactly happened in this one case? | Where did the time go / which hop failed? |
| Shape | Numbers over time, aggregated | Discrete events with detail | One request across every service |
| Cost | Cheap, fixed | Expensive, grows with traffic | Moderate, usually sampled |
| Cardinality | Low — this is the hard constraint | Unlimited | Unlimited |
| Retention | Months to years | Days to weeks | Days |
| Good at | Trends, alerting, SLOs | The specifics of one failure | Latency attribution across services |
| Bad at | Anything per-user | Aggregation | Long-term trend |
Cardinality is the constraint people hit — a metric tagged with customer_id creates one time series per customer, and metrics systems are priced and engineered on the assumption that there are hundreds of series, not millions. High-cardinality dimensions belong in logs and traces — Cardinality.
Traces, and why they’re the one worth adding
In a system with more than one service, a trace is the only signal that reconstructs a request end to end.
trace 7f3a… POST /checkout 1,840ms
├─ auth.verify 12ms
├─ cart.get 34ms
├─ pricing.calculate 88ms
│ └─ tax.lookup 71ms
├─ payment.charge 612ms
└─ inventory.reserve 1,090ms ← here
├─ db.query SELECT … FROM stock 18ms
├─ db.query SELECT … FROM stock 17ms
├─ db.query SELECT … FROM stock 19ms ← ×58, one per line item
└─ …
Two things are visible that no dashboard would have shown: inventory owns 59% of the request, and its cost is an N+1 query rather than one slow query. Neither is a metric anyone would have thought to create.
Propagate the trace context everywhere, including into queues and background jobs, or every asynchronous path appears as an unexplained gap.
Structured logs
Free text is unqueryable. One JSON object per event, with consistent field names, is.
{"level":"error","msg":"payment declined","trace_id":"7f3a…",
"request_id":"req_01H…","order_id":"1234","provider":"stripe",
"decline_code":"insufficient_funds","duration_ms":612}trace_idon every line is what joins logs to traces, and it’s the single highest-value field- Log the identifiers you’ll search by — order, customer, tenant, session
- Never log secrets, card numbers, tokens or full personal records. Logs are widely readable, long-retained and frequently exported. Under UK GDPR — the UK’s retained General Data Protection Regulation — they’re personal data processing like anything else — PII in Analytics, Secrets Management
- One event per line, no multi-line stack traces unless the collector reassembles them
Making it useful in practice
- Instrument the boundaries first — inbound requests, outbound calls, database queries, queue operations. That’s ~80% of the value and most of it is automatic with OpenTelemetry, the vendor-neutral instrumentation standard
- Tag everything with version and deploy ID. “Did this start with the release” is the first question in most incidents, and it’s unanswerable without it — Blue-Green and Rolling Deployments
- Sample intelligently, not uniformly. Keep every error and every slow request; sample the fast successes hard. Uniform sampling throws away exactly the traces you’ll want
- Correlate to the user-visible number. A
request_idreturned in the API response is what connects a customer complaint to a log line — API Design - Watch the bill. Observability spend can rival infrastructure spend, and it grows with traffic — which means it grows fastest during the incident
Where it interacts
- Alerting — what you do with metrics; observability is what you do after the alert fires
- Error Tracking — a specialised slice of logs, grouped and deduplicated, answering a different question
- Service Level Objectives — SLOs are built from metrics, so the metric has to exist and measure the customer’s experience rather than the server’s
- Real User Monitoring — the same idea for the browser half, which server-side observability cannot see at all
- Instrumentation Debugging — the analytics counterpart; the two estates answer different questions and are routinely confused