Tags: web-dev concept

Error Tracking

Date: 2026-08-17


Capturing exceptions, grouping identical ones and ranking what to fix. The grouping is the entire product — a thousand raw stack traces is noise, and the same thousand collapsed into eleven issues with counts and affected-user numbers is a prioritised list.


Error tracking is automatically capturing runtime exceptions with their stack traces and context, grouping occurrences of the same error into issues, and counting them.

What it adds over logs

LOGS                                 ERROR TRACKING

1,240 lines containing "error"       11 issues

grep, then read                      TypeError: cannot read 'sku' of undefined
counts by hand                         cart.js:88 · 940 events · 612 users
no idea if it's new                    first seen 14:02 today, after deploy a3f9
no idea how many users                 ↑ new, regression, top of the list
no stack trace grouping
                                     PaymentDeclined
                                       checkout.js:210 · 180 events · 178 users
                                       first seen 8 months ago, steady
                                       ↑ not a bug. this is business as usual

The second panel is a decision. The first is homework.

Grouping, and its failure modes

Errors are fingerprinted — usually from exception type plus a normalised stack trace — and identical fingerprints collapse into one issue. When it goes wrong it goes wrong in two directions:

  • Over-grouping. Everything lands in one giant Error issue because the message is generic or the frames are all framework code. Fix by throwing typed errors with stable messages and putting the variable part in structured context rather than in the message string — "Payment failed" with {provider, code} attached, never "Payment failed for order 1234", which creates one issue per order
  • Under-grouping. One bug appears as forty issues because minified frames or line numbers shift between releases. Fix with source maps and a stable release identifier — Source Maps

The context that makes an error actionable

A stack trace alone rarely reproduces anything. What’s needed alongside:

  • Release and deploy ID — is this new, and did it start with a release
  • User or session identifier, pseudonymous — how many people, and is it one customer or everyone
  • Breadcrumbs — the last N actions, requests and navigations before the throw. Usually more useful than the trace itself
  • trace_id to jump to the full distributed trace — Observability
  • Environment, browser, device, region — the dimensions that turn “sometimes” into “Safari 17 in the EU”

Triage

The ranking that works, in order:

  1. New since the last deploy — regressions, cheapest to fix while the change is fresh, and the most likely to be your fault
  2. High affected-user count — not high event count. One bot in a retry loop generates 40,000 events and affects nobody
  3. On a revenue path — a checkout error at 60 users beats a settings-page error at 6,000
  4. Rising trend — a slope matters more than a level

Everything else gets ignored or muted, deliberately. The characteristic failure of error tracking is an inbox of 400 permanently unresolved issues, at which point nobody looks at any of them and the tool has become worse than useless — it’s providing the feeling of coverage without the coverage. Muting known-benign issues is discipline, not laziness.

Browser errors are their own problem

The client-side stream is much noisier and a large fraction of it is not yours:

  • Browser extensions throwing inside your page
  • Cross-origin scripts reporting the useless Script error. with no detail — fixed by serving third-party JavaScript with the right CORS headers and crossorigin on the tag — CORS
  • ResizeObserver loop limit exceeded and similar benign browser noise
  • Bots and headless browsers on ancient engines
  • Users on networks that mangle responses

Filter aggressively at the SDK level rather than at the dashboard, or you pay to ingest noise. And without Source Maps uploaded per release, every browser stack trace is minified nonsense — that upload step is the single thing that makes browser error tracking work.

Errors that never throw

The most damaging failures are handled ones. A try/catch that logs and continues shows up nowhere if nothing reports it, and a caught error that silently degrades a feature is invisible by design.

  • Report explicitly from catch blocks where the failure matters, with severity set appropriately
  • Rejected promises without a handler need a global unhandledrejection listener, or async failures vanish
  • Failed network calls returning a fallback are the classic silent one — the page renders, the recommendations are missing, nobody knows

Where it interacts

  • Observability — errors are one signal; the trace is how you find out why
  • Alerting — alert on a new issue affecting many users, or on a rate spike. Never on every error, which is how people learn to ignore the channel
  • Incident Response — an error spike correlated with a deploy is the fastest route to a rollback decision — Rollback and Forward Fix
  • PII in Analytics — breadcrumbs and request bodies routinely capture personal data, including form contents. Scrub at the SDK, before transmission
  • The symptom list — a spike in one error type is often the first visible sign of a tracking or data problem rather than a code one