Tags: web-dev concept

Rate Limiting

Date: 2026-08-17


Capping how much any one caller can consume, so that a system fails a few requests deliberately instead of failing every request accidentally. It’s a capacity control first and an abuse control second — the retry loop that takes you down is usually a customer’s, not an attacker’s.


Rate limiting is restricting the number of requests a caller — a user, key or IP address — may make within a time window, rejecting or delaying the excess.

The algorithms

Four, and the differences matter at the boundaries.

FIXED WINDOW              100/min, counter resets on the minute
  |████████░░|██████░░░░|
   :00      :59 :00
  ✗ 100 requests at 11:59:59 and 100 at 12:00:01 = 200 in two seconds

SLIDING WINDOW            100 in any trailing 60s
  ✓ no boundary burst.  costs more state

TOKEN BUCKET              bucket of 100, refills 10/sec
  ✓ allows a burst up to the bucket size, then settles to the rate
  ✓ the right default — real traffic is bursty and this tolerates it

LEAKY BUCKET              queue drains at a fixed rate
  ✓ perfectly smooth output. adds latency instead of rejecting

Token bucket is the sensible default. A page load firing eight API calls at once is normal behaviour, and an algorithm that rejects it while permitting the same eight spread over a second is punishing correctness.

Token bucket, written out

The whole algorithm is about fifteen lines, and the trick is that nothing runs on a timer — tokens are refilled lazily by multiplying elapsed time by the rate, at the moment someone asks.

const buckets = new Map();   // in production: Redis, so all instances share it
 
function take(key, { capacity = 100, refillPerSec = 10 }) {
  const now = Date.now() / 1000;
  const b = buckets.get(key) ?? { tokens: capacity, last: now };
 
  // refill for however long it's been, never above capacity
  b.tokens = Math.min(capacity, b.tokens + (now - b.last) * refillPerSec);
  b.last = now;
 
  if (b.tokens < 1) {
    buckets.set(key, b);
    // how long until one token exists — this is your Retry-After
    return { ok: false, retryAfter: Math.ceil((1 - b.tokens) / refillPerSec) };
  }
 
  b.tokens -= 1;
  buckets.set(key, b);
  return { ok: true, remaining: Math.floor(b.tokens) };
}

capacity is the burst you tolerate; refillPerSec is the sustained rate you’ll actually allow. A bucket of 100 refilling at 10/sec permits an instant burst of 100, then settles to 10 per second — which is the behaviour real traffic needs and a fixed window can’t express.

Wiring it up, with the response contract from below:

app.use((req, res, next) => {
  const { ok, remaining, retryAfter } = take(req.apiKey ?? req.ip, {
    capacity: 100, refillPerSec: 10,
  });
 
  res.set('RateLimit-Limit', '100');
  res.set('RateLimit-Remaining', String(remaining ?? 0));   // on every response
 
  if (!ok) {
    res.set('Retry-After', String(retryAfter));
    return res.status(429).json({ error: { type: 'rate_limited' } });
  }
  next();
});

The Map is the one thing that doesn’t survive production. With more than one instance, each gets its own bucket and the effective limit multiplies by the instance count — so the counter has to live somewhere shared, and the read-modify-write has to be atomic (a Redis Lua script or INCR with an expiry) rather than three round trips that race — Race Conditions.

What to key on

The identifier decides who gets punished, and getting it wrong is the usual production incident.

KeyGood forFails when
API key / accountThe main case for authenticated APIsDoesn’t protect the login endpoint itself
User IDPer-user fairnessOnly after authentication
IP addressUnauthenticated endpointsCorporate NAT, mobile carriers, schools — thousands of people behind one IP
IP + endpointLogin, password reset, signupSame NAT problem, narrower blast radius
TenantMulti-Tenancy — noisy neighboursNeeds tenant resolved early

IP-based limiting on a consumer site blocks whole offices. It’s still the only option for unauthenticated endpoints, so keep the limits generous, scope them per endpoint, and never apply them to page loads — only to the expensive or abusable ones.

Behind a CDN or proxy, the client IP is in a forwarded header, and that header is user-supplied unless your proxy overwrites it — trusting it naively means every attacker has unlimited IPs.

Responding properly

HTTP/1.1 429 Too Many Requests
Retry-After: 30
RateLimit-Limit: 1000
RateLimit-Remaining: 0
RateLimit-Reset: 1755425430
  • 429, not 503 and not 403. It tells the client this is about rate, is temporary, and is retryable
  • Retry-After always. Without it, a well-behaved client has to guess, and most guess badly
  • Remaining-quota headers on every response, not just rejections, so clients can slow down before being rejected — API Design
  • Never rate-limit silently by slowing responses. It looks like an outage, triggers client timeouts, and holds your own connections open

What actually causes the incidents

Rarely malice:

  • A client’s retry loop with no backoff. One buggy integration produces more load than any attacker bothers to
  • A cron job everyone scheduled at midnight, hitting simultaneously
  • A sync that fetches everything hourly because incremental was harder
  • Your own frontend — a useEffect firing on every keystroke, a polling interval that survived a refactor
  • A crawler doing exactly what crawlers do, across every facet combination — Faceted Navigation and Crawl Budget

Add jitter to anything scheduled or retried, on both sides. Synchronised clients are self-inflicted denial of service — Message Queues.

Where to enforce it

CDN / edge          ← cheapest. rejects before your infrastructure is touched
   ↓
load balancer       ← coarse, effective, no application logic
   ↓
API gateway / BFF   ← backend for frontend. knows the client
                       identity, so this is the usual right place
   ↓
application         ← knows the cost of the operation
   ↓
database            ← last resort: statement timeouts, connection limits

As early as possible, but no earlier than where the identity is known. The edge is cheap but often can’t tell which customer is calling; the application knows everything and has already spent the resources by the time it decides.

Beyond counting requests

  • Cost-based limits for anything where requests aren’t comparable. A GraphQL query can cost anything, so limit computed complexity, not request count — REST GraphQL and RPC
  • Concurrency limits are sometimes the better control — “5 in flight per customer” protects a database connection pool in a way requests-per-minute doesn’t
  • Different limits per operation. Reading a product and generating a report are not the same event
  • Load shedding — under genuine overload, reject cheaply and early rather than accepting work you can’t finish. Shed the least valuable traffic first, which requires having decided in advance what that is — Graceful Degradation
  • Tiered limits as a product feature, where the plan sets the ceiling

Where it interacts

  • Integration Patterns — you are far more often constrained by a vendor’s limits than protected by your own; design syncs around theirs
  • Webhooks — the inbound endpoint needs limiting too, and the sender’s retry storm is exactly the traffic shape that finds the flaw
  • Common Vulnerabilities — rate limiting on login and password reset is a primary defence against credential stuffing, and is separate from capacity protection
  • Bot and Internal Traffic — the same traffic that needs limiting also needs excluding from analytics, and the two systems should agree on what a bot is