Tags: web-dev concept

Error Handling Strategies

Date: 2026-08-17


Fail fast, degrade gracefully, or retry — three strategies for three different kinds of failure. Choosing wrongly is how a transient network blip becomes a lost order, and how a genuine bug gets swallowed and shipped.


Error handling is deciding what happens when something goes wrong. The first question is always the same: is this failure transient or permanent?

TRANSIENT              PERMANENT
network timeout        400 Bad Request
503 Service            401 Unauthorised
  Unavailable          404 Not Found
429 Too Many           422 Unprocessable
  Requests             a bug in your code
deadlock
                       → retrying is
→ retrying may work      pointless and
                         may cause harm

Retrying a permanent failure is the most common error-handling bug, and at scale it’s a self-inflicted denial of service.

The three strategies

FAIL FAST
  stop immediately, surface loudly
  → programmer errors, invalid state,
    misconfiguration
  → a bug caught in development beats
    one absorbed in production

DEGRADE GRACEFULLY
  continue with reduced function
  → non-essential features
  → recommendations fail → hide the
    module, keep the page

RETRY
  try again with backoff
  → transient failures only
  → requires the operation to be
    idempotent

See: Idempotency

Pick per operation, not per application. A payment fails fast; a recommendation carousel degrades; a webhook delivery retries.

Retrying properly

NAIVE                    CORRECT
retry immediately        exponential backoff
retry forever            + jitter
                         + a cap
                         + a total deadline

1s, 2s, 4s, 8s, 16s
  ± random jitter
  max 5 attempts

Jitter is the part that gets omitted and matters most. Without it, every client that failed at the same moment retries at the same moment — the recovering service is hit by a synchronised wave and falls over again.

Retry-After: 120

Honour it when the server sends it — that’s the server telling you exactly what it wants, and ignoring it turns a managed recovery into a fight.

Circuit breakers

When a dependency is down, continuing to call it wastes time and delays its recovery. Deciding what the system does while the circuit is open is Graceful Degradation:

CLOSED     calls pass through
             failures counted
             threshold exceeded ↓
OPEN       calls fail IMMEDIATELY
             no waiting on timeouts
             after a cooldown ↓
HALF-OPEN  let one through
             success → CLOSED
             failure → OPEN

The value is failing fast while broken. Without it, every request waits for a 30-second timeout and your own threads exhaust — which is how one slow dependency takes down a service that doesn’t depend on it much at all.

Timeouts

The control people forget entirely:

NO TIMEOUT   the default in many clients
             → a hung connection holds a
               thread forever

SET ONE      on every network call, always
             shorter than your own caller's
             timeout

Timeouts must decrease as you go deeper. If your API has a 10s budget and calls a service with a 30s timeout, the timeout never fires — your caller gives up first and the work continues pointlessly.

Errors as values versus exceptions

EXCEPTIONS              ERRORS AS VALUES
throw / try / catch     return a result
                        object

invisible in the        visible in the
signature               signature
easy to forget          impossible to ignore
  to catch
clean happy path        verbose

Neither is right in general. The useful convention: exceptions for the exceptional, return values for the expected. A validation failure is an ordinary outcome and should be a value; a database being unreachable is exceptional.

What to do with the error

  • Log with context. “Error” alone is useless; include the operation, the identifiers, and the input that caused it
  • Never swallow silently. catch (e) {} is how bugs live for months. If you genuinely intend to ignore it, log at debug level and comment why
  • Don’t leak internals to users. A stack trace in a response is reconnaissance — Common Vulnerabilities
  • Separate what the user sees from what you record. “Something went wrong, try again” for them; the full detail for you
  • Give the user a way out. An error state with no retry and no alternative is a dead end — Component States
  • Alert on rate, not on instances. One 500 is noise; a rising 500 rate is an incident