Rate Limiting
Date: 2026-08-17
Capping how much any one caller can consume, so that a system fails a few requests deliberately instead of failing every request accidentally. It’s a capacity control first and an abuse control second — the retry loop that takes you down is usually a customer’s, not an attacker’s.
Rate limiting is restricting the number of requests a caller — a user, key or IP address — may make within a time window, rejecting or delaying the excess.
The algorithms
Four, and the differences matter at the boundaries.
FIXED WINDOW 100/min, counter resets on the minute
|████████░░|██████░░░░|
:00 :59 :00
✗ 100 requests at 11:59:59 and 100 at 12:00:01 = 200 in two seconds
SLIDING WINDOW 100 in any trailing 60s
✓ no boundary burst. costs more state
TOKEN BUCKET bucket of 100, refills 10/sec
✓ allows a burst up to the bucket size, then settles to the rate
✓ the right default — real traffic is bursty and this tolerates it
LEAKY BUCKET queue drains at a fixed rate
✓ perfectly smooth output. adds latency instead of rejecting
Token bucket is the sensible default. A page load firing eight API calls at once is normal behaviour, and an algorithm that rejects it while permitting the same eight spread over a second is punishing correctness.
Token bucket, written out
The whole algorithm is about fifteen lines, and the trick is that nothing runs on a timer — tokens are refilled lazily by multiplying elapsed time by the rate, at the moment someone asks.
const buckets = new Map(); // in production: Redis, so all instances share it
function take(key, { capacity = 100, refillPerSec = 10 }) {
const now = Date.now() / 1000;
const b = buckets.get(key) ?? { tokens: capacity, last: now };
// refill for however long it's been, never above capacity
b.tokens = Math.min(capacity, b.tokens + (now - b.last) * refillPerSec);
b.last = now;
if (b.tokens < 1) {
buckets.set(key, b);
// how long until one token exists — this is your Retry-After
return { ok: false, retryAfter: Math.ceil((1 - b.tokens) / refillPerSec) };
}
b.tokens -= 1;
buckets.set(key, b);
return { ok: true, remaining: Math.floor(b.tokens) };
}capacity is the burst you tolerate; refillPerSec is the sustained rate you’ll actually allow. A bucket of 100 refilling at 10/sec permits an instant burst of 100, then settles to 10 per second — which is the behaviour real traffic needs and a fixed window can’t express.
Wiring it up, with the response contract from below:
app.use((req, res, next) => {
const { ok, remaining, retryAfter } = take(req.apiKey ?? req.ip, {
capacity: 100, refillPerSec: 10,
});
res.set('RateLimit-Limit', '100');
res.set('RateLimit-Remaining', String(remaining ?? 0)); // on every response
if (!ok) {
res.set('Retry-After', String(retryAfter));
return res.status(429).json({ error: { type: 'rate_limited' } });
}
next();
});The Map is the one thing that doesn’t survive production. With more than one instance, each gets its own bucket and the effective limit multiplies by the instance count — so the counter has to live somewhere shared, and the read-modify-write has to be atomic (a Redis Lua script or INCR with an expiry) rather than three round trips that race — Race Conditions.
What to key on
The identifier decides who gets punished, and getting it wrong is the usual production incident.
| Key | Good for | Fails when |
|---|---|---|
| API key / account | The main case for authenticated APIs | Doesn’t protect the login endpoint itself |
| User ID | Per-user fairness | Only after authentication |
| IP address | Unauthenticated endpoints | Corporate NAT, mobile carriers, schools — thousands of people behind one IP |
| IP + endpoint | Login, password reset, signup | Same NAT problem, narrower blast radius |
| Tenant | Multi-Tenancy — noisy neighbours | Needs tenant resolved early |
IP-based limiting on a consumer site blocks whole offices. It’s still the only option for unauthenticated endpoints, so keep the limits generous, scope them per endpoint, and never apply them to page loads — only to the expensive or abusable ones.
Behind a CDN or proxy, the client IP is in a forwarded header, and that header is user-supplied unless your proxy overwrites it — trusting it naively means every attacker has unlimited IPs.
Responding properly
HTTP/1.1 429 Too Many Requests
Retry-After: 30
RateLimit-Limit: 1000
RateLimit-Remaining: 0
RateLimit-Reset: 1755425430
429, not503and not403. It tells the client this is about rate, is temporary, and is retryableRetry-Afteralways. Without it, a well-behaved client has to guess, and most guess badly- Remaining-quota headers on every response, not just rejections, so clients can slow down before being rejected — API Design
- Never rate-limit silently by slowing responses. It looks like an outage, triggers client timeouts, and holds your own connections open
What actually causes the incidents
Rarely malice:
- A client’s retry loop with no backoff. One buggy integration produces more load than any attacker bothers to
- A cron job everyone scheduled at midnight, hitting simultaneously
- A sync that fetches everything hourly because incremental was harder
- Your own frontend — a
useEffectfiring on every keystroke, a polling interval that survived a refactor - A crawler doing exactly what crawlers do, across every facet combination — Faceted Navigation and Crawl Budget
Add jitter to anything scheduled or retried, on both sides. Synchronised clients are self-inflicted denial of service — Message Queues.
Where to enforce it
CDN / edge ← cheapest. rejects before your infrastructure is touched
↓
load balancer ← coarse, effective, no application logic
↓
API gateway / BFF ← backend for frontend. knows the client
identity, so this is the usual right place
↓
application ← knows the cost of the operation
↓
database ← last resort: statement timeouts, connection limits
As early as possible, but no earlier than where the identity is known. The edge is cheap but often can’t tell which customer is calling; the application knows everything and has already spent the resources by the time it decides.
Beyond counting requests
- Cost-based limits for anything where requests aren’t comparable. A GraphQL query can cost anything, so limit computed complexity, not request count — REST GraphQL and RPC
- Concurrency limits are sometimes the better control — “5 in flight per customer” protects a database connection pool in a way requests-per-minute doesn’t
- Different limits per operation. Reading a product and generating a report are not the same event
- Load shedding — under genuine overload, reject cheaply and early rather than accepting work you can’t finish. Shed the least valuable traffic first, which requires having decided in advance what that is — Graceful Degradation
- Tiered limits as a product feature, where the plan sets the ceiling
Where it interacts
- Integration Patterns — you are far more often constrained by a vendor’s limits than protected by your own; design syncs around theirs
- Webhooks — the inbound endpoint needs limiting too, and the sender’s retry storm is exactly the traffic shape that finds the flaw
- Common Vulnerabilities — rate limiting on login and password reset is a primary defence against credential stuffing, and is separate from capacity protection
- Bot and Internal Traffic — the same traffic that needs limiting also needs excluding from analytics, and the two systems should agree on what a bot is