Tags: web-dev concept

Graceful Degradation

Date: 2026-08-17


Deciding in advance which features stop working first, so a partial failure stays partial. Without the decision, every dependency is implicitly critical — and a review service being slow takes down checkout, which nobody would have chosen if asked.


Graceful degradation is designing a system so that when a dependency fails, it keeps serving with reduced functionality instead of failing entirely.

The decision, made once per dependency

product page dependencies       if it fails, the page should…

catalogue        CRITICAL       fail. there is no page without it
price            CRITICAL       fail. showing a wrong or missing price is worse
                                than showing nothing
stock            DEGRADE        render, hide the availability badge,
                                let checkout re-check
reviews          DEGRADE        render without the reviews block
recommendations  DEGRADE        render without the carousel
personalisation  DEGRADE        render the default
analytics        IGNORE         never block a render on measurement

That table is the entire concept. Writing it down is most of the work, and the argument it settles — is stock critical? — is the valuable part. Undecided, the answer is decided by an exception handler nobody looked at, and it’s usually “critical”.

Timeouts are the mechanism

A dependency without a timeout is not degradable, because “failing” and “taking 30 seconds” are the same thing to a customer and only the first one triggers your fallback.

no timeout          reviews hang → request thread held → pool exhausted
                    → the whole site is down because of a reviews outage

timeout 300ms       reviews hang → abandoned at 300ms → page renders without them
                    → customers never notice

Every network call gets a timeout, and the timeout is a product decision. Set it from what the page can afford, not from what the dependency usually takes — a 250ms budget for a non-critical widget on a page targeting a 2.5s Largest Contentful Paint.

Written out, the whole degradation policy for one dependency is a wrapper:

async function degradable(name, fn, { timeoutMs, fallback }) {
  const ac = new AbortController();
  const timer = setTimeout(() => ac.abort(), timeoutMs);
  try {
    return await fn(ac.signal);
  } catch (err) {
    // degrading is a normal outcome, but it must never be silent
    metrics.increment('dependency.degraded', { name });
    return fallback;
  } finally {
    clearTimeout(timer);
  }
}
 
// the table above, as code — each line is one row of it
const [price, reviews, recs] = await Promise.all([
  pricing.get(sku),                                    // CRITICAL: no wrapper.
                                                       // it throws, the page 500s, correctly
  degradable('reviews', s => fetchReviews(sku, s),
    { timeoutMs: 250, fallback: null }),               // null → block doesn't render
  degradable('recs', s => fetchRecs(sku, s),
    { timeoutMs: 150, fallback: BESTSELLERS }),        // static default, nobody notices
]);

Two things the code makes obvious that prose doesn’t. Promise.all means the timeouts run concurrently, so the page’s worst case is the longest single timeout rather than their sum. And the critical dependency is marked by the absence of a wrapper — which is why the decision has to be written down somewhere, or “no wrapper” reads as “not got round to it yet”.

Circuit breakers

Timeouts alone still spend the full timeout on every request while a dependency is down — and worse, keep hammering something that’s trying to recover.

CLOSED      normal. calls pass through, failures counted
              │  failure rate over threshold
              ▼
OPEN        calls fail instantly, no request made
            the dependency gets breathing room, you stop wasting time
              │  after a cool-off
              ▼
HALF-OPEN   let a few through
              ├─ they succeed → CLOSED
              └─ they fail    → OPEN again

The second-order benefit matters as much as the first: a struggling service that receives every retry never recovers, and an open circuit is what lets it come back.

Designing the fallback

The fallback is a product decision, not an empty catch.

FallbackWhen it fits
Stale cacheBest available. Yesterday’s recommendations beat none — Caching Strategies
Static defaultGeneric bestsellers instead of personalised — nobody can tell
Hide the componentReviews, badges. Layout must not shift when it vanishes — Cumulative Layout Shift
Queue itWrites that can complete later — Message Queues
Reduced functionBrowse-only mode with checkout disabled and said so
Honest errorOnly where the alternative is misleading — payment, stock at the point of commitment

Never a fallback that lies about money or availability. Showing a stale price or “in stock” from a cache during an outage converts a technical problem into a customer-service and possibly a consumer-law one.

Test the fallback path. Untested fallbacks are broken at roughly the rate of any other untested code, and they only run on the worst day. Force failures deliberately — a flag that disables a dependency, exercised in production during quiet hours.

Prioritising under load

Degradation isn’t only about dependencies failing — it’s also what to abandon when you can’t serve everything.

  • Shed the cheapest-to-lose traffic first: personalisation, recommendations, non-essential API consumers, then reporting. Checkout last
  • Protect writes over reads. A customer mid-purchase matters more than one browsing
  • Serve stale rather than nothing. stale-while-revalidate at the CDN keeps a site up through a total origin outage — CDN Caching, HTTP Caching
  • A static fallback page at the edge, deployed and tested, is the last line and is worth having

Not the same as progressive enhancement

Related and routinely conflated:

  • Progressive Enhancement starts from a working baseline and layers capability on — a browser-capability strategy, decided at build time
  • Graceful degradation starts from a full experience and defines what’s shed — a runtime failure strategy for services

They meet in one place: a page that renders server-side and works without JavaScript degrades gracefully by construction, because the client-side dependencies were never load-bearing — Rendering Strategies.

Where it interacts

  • Backend for Frontend — the natural place to implement per-client degradation policy, since it’s the only layer that knows which client is asking
  • Service Level Objectives — degraded is neither up nor down, so the service level indicator (SLI) has to define which side of the line it falls on
  • Monolith vs Services — fault isolation is a claimed benefit of services and only materialises if this work is done
  • Feature Flags — the kill switch that makes degradation available as a manual option during an incident — Incident Response