Message Queues
Date: 2026-08-17
A buffer between a producer and a consumer, so the producer doesn’t wait and doesn’t care whether the consumer is up. It converts a synchronous failure into a delayed success — which is the entire benefit, and also the reason it’s the wrong tool whenever the caller genuinely needs the answer now.
A message queue is a durable buffer that producers write messages to and consumers read from independently, so the two don’t need to be running, or fast, at the same time.
What it changes
SYNCHRONOUS QUEUED
checkout checkout
→ charge card 800ms → charge card 800ms
→ write order 40ms → write order 40ms
→ email receipt 1200ms ✗ → enqueue receipt job 2ms
→ sync to warehouse 900ms ✗ → enqueue warehouse job 2ms
→ update analytics 300ms ✗ → 200 OK ~850ms
→ 200 OK 3240ms
workers drain the queue at their own pace
if the email provider is down, if the email provider is down, the job
the customer's order fails retries. the order was never at risk
The customer’s request now depends only on the things that must be true before responding. Everything else moved behind a queue where failure is recoverable.
The two shapes
| Queue (work distribution) | Topic / log (publish-subscribe) | |
|---|---|---|
| Consumers | Compete — each message goes to one | Independent — each gets every message |
| Message after consumption | Deleted | Retained; each consumer tracks its own position |
| Adding a consumer | Splits the work | Adds a new reader of the same stream |
| Typical use | Send the email, resize the image | Order placed → warehouse, CRM, analytics all react |
| Examples | SQS, RabbitMQ queues, Sidekiq | Kafka, Kinesis, Pub/Sub |
The log shape is what makes Event-Driven Architecture practical: a new consumer can be added later and replay from the beginning, which a delete-on-read queue can’t offer.
Delivery guarantees
The three options, and only two of them exist.
at-most-once ack before processing. crash → message lost. rarely acceptable
at-least-once ack after processing. crash → message redelivered. THE DEFAULT
exactly-once marketing, mostly
“Exactly-once delivery” is not achievable across a network. What’s achievable is at-least-once delivery plus idempotent processing, which produces exactly-once effect. Where a broker advertises exactly-once, it means within its own transactional boundary — the moment your consumer calls an external API, you’re back to needing idempotency — Idempotency.
So: every consumer must be safe to run twice on the same message. Deduplicate on a message ID, or make the operation naturally idempotent (SET status = 'shipped' rather than INCREMENT attempts).
The safe consumer, written out:
async function handle(message) {
await db.transaction(async (tx) => {
// insert-first: the unique constraint IS the deduplication.
// checking with a SELECT first would race two workers against each other
const [claimed] = await tx('processed_messages')
.insert({ message_id: message.id })
.onConflict('message_id').ignore()
.returning('message_id');
if (!claimed) return; // duplicate. do nothing, still ack below
await tx('orders')
.where({ id: message.orderId })
.update({ status: 'shipped' }); // the effect and the claim commit together
});
await ack(message); // ack AFTER the commit — at-least-once
}Three things are load-bearing. The claim and the effect share one transaction, so a crash between them rolls back both and the redelivery works. The insert is the check — a SELECT followed by an INSERT is a race two competing workers will find. And the ack comes last: acking first turns this into at-most-once and loses the message on any failure.
Failure handling
- Retry with exponential backoff and jitter. Immediate retries turn a struggling downstream service into a dead one; retries without jitter make every worker retry in unison
- A dead-letter queue. After N attempts the message moves aside rather than blocking or vanishing. A dead-letter queue nobody looks at is a folder of lost orders — alert on its depth, and have a documented replay path
- Poison messages. One malformed message that crashes the consumer on every attempt will, in a strictly-ordered queue, stop everything behind it. This is the main hazard of ordering guarantees
- Visibility timeout. If it’s shorter than the work takes, the message is redelivered while still being processed — a duplicate you caused yourself
The numbers that matter
Queue depth and consumer lag are the two health metrics, and they say different things.
depth rising, lag rising consumers can't keep up — scale out
depth rising, lag flat producer spike, consumers coping
depth zero, lag zero healthy, or nothing is being produced ← alert on this too
oldest message age the honest one: how stale is the worst case
Alert on oldest-message-age rather than depth. Depth of 10,000 draining in 30 seconds is fine; depth of 3 stuck for an hour is an incident — Alerting.
Ordering
Global ordering and parallel consumers are mutually exclusive. What you get instead is ordering within a partition key — all messages for order 1234 in sequence, different orders in parallel. Choose the key carefully: too coarse and you’ve serialised everything, too fine and related messages race.
When not to use one
- The caller needs the result. A queue plus polling for a response is a slow, complicated remote procedure call — REST GraphQL and RPC
- The work is fast and the failure is the caller’s problem. Validating a postcode doesn’t need a queue
- You have one server and low volume. A database table with a
statuscolumn and a cron worker is a queue, is debuggable with SQL, and is often the right amount of machinery - It’s hiding a slow dependency you could fix. Queues make latency invisible rather than absent
Where it interacts
- Webhooks — the receive-then-enqueue pattern is what makes webhook endpoints survivable
- Race Conditions — queues remove some by serialising per key, and introduce others through out-of-order processing
- Eventual Consistency — anything behind a queue is stale by definition, so the UI has to be honest about it
- Observability — propagate the trace context into the message, or every async path becomes an unexplained gap