Webhooks
Date: 2026-08-17
A webhook is an HTTP request someone else’s system makes to yours when something happens there. It replaces polling, and in exchange hands you every problem of running a public endpoint that must never lose a message, never trust its caller, and never assume anything arrives once or in order.
What changes versus polling
POLLING WEBHOOK
you → GET /orders?since=… them → POST https://you/hooks/orders
← nothing ← 200
every 60s, mostly nothing fires when it happens
you control the rate they control the rate
delay = poll interval delay = milliseconds
their outage is invisible your outage loses events
you now run a public endpoint
You’ve traded a known cost for a set of unknown failure modes. Polling is unfashionable and still correct when volume is low and latency doesn’t matter.
The five rules
1. Verify the signature. Your endpoint is a public URL that performs privileged actions on unauthenticated input. Providers sign the raw body with a shared secret; you recompute and compare.
const expected = crypto
.createHmac('sha256', process.env.WEBHOOK_SECRET)
.update(rawBody) // raw bytes — parsed-then-restringified won't match
.digest('hex');
if (!crypto.timingSafeEqual(Buffer.from(expected), Buffer.from(header))) {
return res.status(401).end();
}Two things people get wrong: signing the re-serialised JSON rather than the raw body, and comparing with ===, which leaks the secret through timing — Hashing, Common Vulnerabilities.
Check the timestamp too, and reject anything older than a few minutes. Without that, a captured request can be replayed forever.
2. Respond immediately, process later. The sender is holding a connection open with a short timeout.
receive → verify signature → write to a queue → 200 OK ~15ms
↓
worker processes it, retries on failure
Doing the work inline means a slow database turns into a timeout, which turns into the provider retrying, which turns into duplicate processing under load — the exact moment you least want it — Message Queues.
3. Assume duplicates. Delivery is at-least-once: the sender retries when it doesn’t get a 200, including when it did get one and the response was lost. Every handler must be safe to run twice, keyed on the provider’s event ID — Idempotency.
4. Assume out-of-order. order.updated can arrive before order.created. Order by the event’s own timestamp or version field, not by arrival — and ignore an event older than the state you already hold.
5. Treat the payload as a notification, not as truth. It may be stale by the time you process it, and it’s the thin end of a system you don’t control. For anything that matters — money, stock — fetch the current state from their API using the ID in the payload.
What still goes wrong
- You were down for an hour. Most providers retry with backoff for a while, then give up. Design for a reconciliation job that fetches everything changed since the last known good timestamp — this is the backstop that makes the whole pattern survivable, and it’s the piece almost nobody builds until after the first incident
- Silent stops. Nothing arriving looks identical to nothing happening. Alert on absence — no events in a window that normally has them — Alerting
- Retry storms. Returning 5xx to a high-volume sender queues a retry for every message, and the recovery traffic arrives all at once
- Ordering across entities. Even per-entity ordering doesn’t give you cross-entity ordering, so don’t write logic that assumes it
- Endpoint enumeration. The URL leaks eventually. Signature verification is the control, not obscurity
Sending them
Everything above, inverted:
- Sign the raw body, publish the algorithm, and support two live secrets so consumers can rotate without downtime
- Retry with exponential backoff and jitter, then dead-letter. Give consumers a way to replay from the dead-letter queue
- Include an event ID, an event type, a timestamp and a version in every payload. Without an ID, consumers cannot deduplicate
- Give them a log. “Which of my events failed and why” is the first thing every integrator asks for
- Disable endpoints that fail persistently and notify the owner, or a dead consumer becomes your outbound traffic problem
Where it interacts
- Integration Patterns — webhooks are the point-to-point option, and the mess they accumulate is the reason the other options exist
- Event-Driven Architecture — webhooks are event-driven integration across an organisational boundary, with none of the delivery guarantees an internal broker gives you
- Composable Commerce — a composed stack is largely held together by webhooks, so the reconciliation job stops being optional