Graceful Degradation
Date: 2026-08-17
Deciding in advance which features stop working first, so a partial failure stays partial. Without the decision, every dependency is implicitly critical — and a review service being slow takes down checkout, which nobody would have chosen if asked.
Graceful degradation is designing a system so that when a dependency fails, it keeps serving with reduced functionality instead of failing entirely.
The decision, made once per dependency
product page dependencies if it fails, the page should…
catalogue CRITICAL fail. there is no page without it
price CRITICAL fail. showing a wrong or missing price is worse
than showing nothing
stock DEGRADE render, hide the availability badge,
let checkout re-check
reviews DEGRADE render without the reviews block
recommendations DEGRADE render without the carousel
personalisation DEGRADE render the default
analytics IGNORE never block a render on measurement
That table is the entire concept. Writing it down is most of the work, and the argument it settles — is stock critical? — is the valuable part. Undecided, the answer is decided by an exception handler nobody looked at, and it’s usually “critical”.
Timeouts are the mechanism
A dependency without a timeout is not degradable, because “failing” and “taking 30 seconds” are the same thing to a customer and only the first one triggers your fallback.
no timeout reviews hang → request thread held → pool exhausted
→ the whole site is down because of a reviews outage
timeout 300ms reviews hang → abandoned at 300ms → page renders without them
→ customers never notice
Every network call gets a timeout, and the timeout is a product decision. Set it from what the page can afford, not from what the dependency usually takes — a 250ms budget for a non-critical widget on a page targeting a 2.5s Largest Contentful Paint.
Written out, the whole degradation policy for one dependency is a wrapper:
async function degradable(name, fn, { timeoutMs, fallback }) {
const ac = new AbortController();
const timer = setTimeout(() => ac.abort(), timeoutMs);
try {
return await fn(ac.signal);
} catch (err) {
// degrading is a normal outcome, but it must never be silent
metrics.increment('dependency.degraded', { name });
return fallback;
} finally {
clearTimeout(timer);
}
}
// the table above, as code — each line is one row of it
const [price, reviews, recs] = await Promise.all([
pricing.get(sku), // CRITICAL: no wrapper.
// it throws, the page 500s, correctly
degradable('reviews', s => fetchReviews(sku, s),
{ timeoutMs: 250, fallback: null }), // null → block doesn't render
degradable('recs', s => fetchRecs(sku, s),
{ timeoutMs: 150, fallback: BESTSELLERS }), // static default, nobody notices
]);Two things the code makes obvious that prose doesn’t. Promise.all means the timeouts run concurrently, so the page’s worst case is the longest single timeout rather than their sum. And the critical dependency is marked by the absence of a wrapper — which is why the decision has to be written down somewhere, or “no wrapper” reads as “not got round to it yet”.
Circuit breakers
Timeouts alone still spend the full timeout on every request while a dependency is down — and worse, keep hammering something that’s trying to recover.
CLOSED normal. calls pass through, failures counted
│ failure rate over threshold
▼
OPEN calls fail instantly, no request made
the dependency gets breathing room, you stop wasting time
│ after a cool-off
▼
HALF-OPEN let a few through
├─ they succeed → CLOSED
└─ they fail → OPEN again
The second-order benefit matters as much as the first: a struggling service that receives every retry never recovers, and an open circuit is what lets it come back.
Designing the fallback
The fallback is a product decision, not an empty catch.
| Fallback | When it fits |
|---|---|
| Stale cache | Best available. Yesterday’s recommendations beat none — Caching Strategies |
| Static default | Generic bestsellers instead of personalised — nobody can tell |
| Hide the component | Reviews, badges. Layout must not shift when it vanishes — Cumulative Layout Shift |
| Queue it | Writes that can complete later — Message Queues |
| Reduced function | Browse-only mode with checkout disabled and said so |
| Honest error | Only where the alternative is misleading — payment, stock at the point of commitment |
Never a fallback that lies about money or availability. Showing a stale price or “in stock” from a cache during an outage converts a technical problem into a customer-service and possibly a consumer-law one.
Test the fallback path. Untested fallbacks are broken at roughly the rate of any other untested code, and they only run on the worst day. Force failures deliberately — a flag that disables a dependency, exercised in production during quiet hours.
Prioritising under load
Degradation isn’t only about dependencies failing — it’s also what to abandon when you can’t serve everything.
- Shed the cheapest-to-lose traffic first: personalisation, recommendations, non-essential API consumers, then reporting. Checkout last
- Protect writes over reads. A customer mid-purchase matters more than one browsing
- Serve stale rather than nothing.
stale-while-revalidateat the CDN keeps a site up through a total origin outage — CDN Caching, HTTP Caching - A static fallback page at the edge, deployed and tested, is the last line and is worth having
Not the same as progressive enhancement
Related and routinely conflated:
- Progressive Enhancement starts from a working baseline and layers capability on — a browser-capability strategy, decided at build time
- Graceful degradation starts from a full experience and defines what’s shed — a runtime failure strategy for services
They meet in one place: a page that renders server-side and works without JavaScript degrades gracefully by construction, because the client-side dependencies were never load-bearing — Rendering Strategies.
Where it interacts
- Backend for Frontend — the natural place to implement per-client degradation policy, since it’s the only layer that knows which client is asking
- Service Level Objectives — degraded is neither up nor down, so the service level indicator (SLI) has to define which side of the line it falls on
- Monolith vs Services — fault isolation is a claimed benefit of services and only materialises if this work is done
- Feature Flags — the kill switch that makes degradation available as a manual option during an incident — Incident Response