Error Handling Strategies
Date: 2026-08-17
Fail fast, degrade gracefully, or retry — three strategies for three different kinds of failure. Choosing wrongly is how a transient network blip becomes a lost order, and how a genuine bug gets swallowed and shipped.
Error handling is deciding what happens when something goes wrong. The first question is always the same: is this failure transient or permanent?
TRANSIENT PERMANENT
network timeout 400 Bad Request
503 Service 401 Unauthorised
Unavailable 404 Not Found
429 Too Many 422 Unprocessable
Requests a bug in your code
deadlock
→ retrying is
→ retrying may work pointless and
may cause harm
Retrying a permanent failure is the most common error-handling bug, and at scale it’s a self-inflicted denial of service.
The three strategies
FAIL FAST
stop immediately, surface loudly
→ programmer errors, invalid state,
misconfiguration
→ a bug caught in development beats
one absorbed in production
DEGRADE GRACEFULLY
continue with reduced function
→ non-essential features
→ recommendations fail → hide the
module, keep the page
RETRY
try again with backoff
→ transient failures only
→ requires the operation to be
idempotent
See: Idempotency
Pick per operation, not per application. A payment fails fast; a recommendation carousel degrades; a webhook delivery retries.
Retrying properly
NAIVE CORRECT
retry immediately exponential backoff
retry forever + jitter
+ a cap
+ a total deadline
1s, 2s, 4s, 8s, 16s
± random jitter
max 5 attempts
Jitter is the part that gets omitted and matters most. Without it, every client that failed at the same moment retries at the same moment — the recovering service is hit by a synchronised wave and falls over again.
Retry-After: 120
Honour it when the server sends it — that’s the server telling you exactly what it wants, and ignoring it turns a managed recovery into a fight.
Circuit breakers
When a dependency is down, continuing to call it wastes time and delays its recovery. Deciding what the system does while the circuit is open is Graceful Degradation:
CLOSED calls pass through
failures counted
threshold exceeded ↓
OPEN calls fail IMMEDIATELY
no waiting on timeouts
after a cooldown ↓
HALF-OPEN let one through
success → CLOSED
failure → OPEN
The value is failing fast while broken. Without it, every request waits for a 30-second timeout and your own threads exhaust — which is how one slow dependency takes down a service that doesn’t depend on it much at all.
Timeouts
The control people forget entirely:
NO TIMEOUT the default in many clients
→ a hung connection holds a
thread forever
SET ONE on every network call, always
shorter than your own caller's
timeout
Timeouts must decrease as you go deeper. If your API has a 10s budget and calls a service with a 30s timeout, the timeout never fires — your caller gives up first and the work continues pointlessly.
Errors as values versus exceptions
EXCEPTIONS ERRORS AS VALUES
throw / try / catch return a result
object
invisible in the visible in the
signature signature
easy to forget impossible to ignore
to catch
clean happy path verbose
Neither is right in general. The useful convention: exceptions for the exceptional, return values for the expected. A validation failure is an ordinary outcome and should be a value; a database being unreachable is exceptional.
What to do with the error
- Log with context. “Error” alone is useless; include the operation, the identifiers, and the input that caused it
- Never swallow silently.
catch (e) {}is how bugs live for months. If you genuinely intend to ignore it, log at debug level and comment why - Don’t leak internals to users. A stack trace in a response is reconnaissance — Common Vulnerabilities
- Separate what the user sees from what you record. “Something went wrong, try again” for them; the full detail for you
- Give the user a way out. An error state with no retry and no alternative is a dead end — Component States
- Alert on rate, not on instances. One 500 is noise; a rising 500 rate is an incident