Tags: web-dev concept

Incident Response

Date: 2026-08-17


A rehearsed sequence for when production is broken, and a written account afterwards of why it was possible. The response limits the damage; the account is the only part that compounds — an incident you don’t learn from is an incident you’ve bought twice.


Incident response is the defined process for detecting, coordinating, mitigating and communicating a production failure; a post-incident review (postmortem) is the written analysis that follows.

The sequence

1  DECLARE      say the word "incident" out loud. open a channel.
                under-declaring is the common error and it costs nothing
                to stand one down

2  ROLES        incident commander — coordinates, decides, doesn't debug
                communicator     — updates stakeholders and the status page
                operators        — actually investigate
                (one person can hold several roles on a small team;
                 the commander should not also be debugging)

3  MITIGATE     stop the bleeding. NOT find the cause.
                roll back, flip the flag, fail over, disable the feature

4  COMMUNICATE  first update within minutes, then on a fixed cadence
                even when there's nothing new. silence is read as chaos

5  RESOLVE      confirm recovery with the metric that showed the problem,
                not by looking at the site once

6  REVIEW       blameless postmortem, written, within a few days

Step 3 is the one people get wrong. The instinct is to understand the problem before acting, and understanding takes an hour while a rollback takes ninety seconds. Diagnosis is what happens after customers are served again — Rollback and Forward Fix.

The incident commander

The role that makes the difference on a team of more than three, and it isn’t a technical role.

The commander holds the state, decides, and protects the operators — running the timeline, asking “what would fix this now”, assigning one owner per line of investigation, keeping a lid on the number of people typing, and being the single interface for everyone asking “any update”.

The most common failure is the most senior engineer taking the role and then debugging anyway. Once the commander’s head is inside the problem, nobody is running the incident: three people investigate the same theory, nobody has told the business anything for forty minutes, and the decision to roll back gets made an hour late.

Communicating during

  • First update fast, even with nothing to say. “We’re aware of checkout failures, investigating, next update in 15 minutes” is the whole job
  • Fixed cadence, held. Predictability is what stops people asking, which is what lets the operators work
  • Impact in customer terms, not system terms. “Customers can’t complete checkout” — not “the order service is returning 503s”
  • Never speculate on cause or on time to fix, publicly. Both will be wrong and both get quoted back
  • Say when it’s over, and to everyone you told it had started

The postmortem

Written for every meaningful incident, within a few days while memory is intact.

  • A timeline with timestamps — when it started, when it was detected, when it was declared, when mitigated, when resolved. The gap between started and detected is usually the most useful number in the document, and it’s a direct critique of the alerting
  • Impact, quantified — customers affected, orders lost, revenue, duration. Necessary for prioritising the actions
  • Contributing factors, plural. “The cause” is almost always several things that individually would have been survivable
  • What went well, genuinely — the rollback that worked, the alert that fired correctly. It tells you what to protect
  • Action items with owners and dates. A postmortem with no owned actions is a diary entry

Blameless is a technique, not kindness. People who expect blame withhold detail, and the detail is the entire value. The premise is that anyone in that position with that information would plausibly have done the same, so the question is never “why did you deploy it” but “why was it possible to deploy it, and why didn’t anything catch it”. Systems are what you can change; individual carefulness is not.

Five whys, honestly applied, tends to walk from human error towards the system:

checkout was down 40 minutes
 ← the deploy contained a broken migration
 ← the migration wasn't tested against production-scale data
 ← staging has 12,000 rows and production has 14 million
 ← there is no production-scale environment                    ← fundable
 ← nobody has owned test data since the platform team split    ← the real one

Stopping at “the engineer should have checked” produces no action anyone can take. Two more steps produce a budget line — Environments.

What makes response fast

Built before the incident, never during:

  • Runbooks per alert — what it means, how to confirm, first three things to try — Alerting
  • A rollback that’s one command and has been rehearsed this quarter
  • Kill switches for anything expensive or third-party, so degrading is available as an option — Feature Flags, Graceful Degradation
  • Deploy markers on every dashboard. “Did this start with a release” is the first question every time — Observability
  • A named on-call rotation with escalation, and access rights that don’t need a request at 3am
  • Practice. Game days, or restoring a backup on a quiet Tuesday. An untested recovery path is a hypothesis

Where it interacts

  • Service Level Objectives — the error budget is what an incident spends, which makes severity and follow-up investment arguable in numbers
  • Rollback and Forward Fix — the mitigation decision, which should already have been made in principle
  • Error Tracking — usually the fastest route from “something’s wrong” to “which release”
  • Technical Debt — postmortem actions are the most fundable debt repayment available, because the interest was just paid publicly