Tags: web-dev concept

Message Queues

Date: 2026-08-17


A buffer between a producer and a consumer, so the producer doesn’t wait and doesn’t care whether the consumer is up. It converts a synchronous failure into a delayed success — which is the entire benefit, and also the reason it’s the wrong tool whenever the caller genuinely needs the answer now.


A message queue is a durable buffer that producers write messages to and consumers read from independently, so the two don’t need to be running, or fast, at the same time.

What it changes

SYNCHRONOUS                          QUEUED

checkout                             checkout
  → charge card         800ms          → charge card              800ms
  → write order          40ms          → write order               40ms
  → email receipt      1200ms  ✗       → enqueue receipt job        2ms
  → sync to warehouse   900ms  ✗       → enqueue warehouse job      2ms
  → update analytics    300ms  ✗       → 200 OK                   ~850ms
  → 200 OK             3240ms
                                     workers drain the queue at their own pace
if the email provider is down,       if the email provider is down, the job
the customer's order fails           retries. the order was never at risk

The customer’s request now depends only on the things that must be true before responding. Everything else moved behind a queue where failure is recoverable.

The two shapes

Queue (work distribution)Topic / log (publish-subscribe)
ConsumersCompete — each message goes to oneIndependent — each gets every message
Message after consumptionDeletedRetained; each consumer tracks its own position
Adding a consumerSplits the workAdds a new reader of the same stream
Typical useSend the email, resize the imageOrder placed → warehouse, CRM, analytics all react
ExamplesSQS, RabbitMQ queues, SidekiqKafka, Kinesis, Pub/Sub

The log shape is what makes Event-Driven Architecture practical: a new consumer can be added later and replay from the beginning, which a delete-on-read queue can’t offer.

Delivery guarantees

The three options, and only two of them exist.

at-most-once    ack before processing.  crash → message lost.       rarely acceptable
at-least-once   ack after processing.   crash → message redelivered. THE DEFAULT
exactly-once    marketing, mostly

“Exactly-once delivery” is not achievable across a network. What’s achievable is at-least-once delivery plus idempotent processing, which produces exactly-once effect. Where a broker advertises exactly-once, it means within its own transactional boundary — the moment your consumer calls an external API, you’re back to needing idempotency — Idempotency.

So: every consumer must be safe to run twice on the same message. Deduplicate on a message ID, or make the operation naturally idempotent (SET status = 'shipped' rather than INCREMENT attempts).

The safe consumer, written out:

async function handle(message) {
  await db.transaction(async (tx) => {
    // insert-first: the unique constraint IS the deduplication.
    // checking with a SELECT first would race two workers against each other
    const [claimed] = await tx('processed_messages')
      .insert({ message_id: message.id })
      .onConflict('message_id').ignore()
      .returning('message_id');
 
    if (!claimed) return;                    // duplicate. do nothing, still ack below
 
    await tx('orders')
      .where({ id: message.orderId })
      .update({ status: 'shipped' });        // the effect and the claim commit together
  });
 
  await ack(message);                        // ack AFTER the commit — at-least-once
}

Three things are load-bearing. The claim and the effect share one transaction, so a crash between them rolls back both and the redelivery works. The insert is the check — a SELECT followed by an INSERT is a race two competing workers will find. And the ack comes last: acking first turns this into at-most-once and loses the message on any failure.

Failure handling

  • Retry with exponential backoff and jitter. Immediate retries turn a struggling downstream service into a dead one; retries without jitter make every worker retry in unison
  • A dead-letter queue. After N attempts the message moves aside rather than blocking or vanishing. A dead-letter queue nobody looks at is a folder of lost orders — alert on its depth, and have a documented replay path
  • Poison messages. One malformed message that crashes the consumer on every attempt will, in a strictly-ordered queue, stop everything behind it. This is the main hazard of ordering guarantees
  • Visibility timeout. If it’s shorter than the work takes, the message is redelivered while still being processed — a duplicate you caused yourself

The numbers that matter

Queue depth and consumer lag are the two health metrics, and they say different things.

depth rising, lag rising     consumers can't keep up — scale out
depth rising, lag flat       producer spike, consumers coping
depth zero, lag zero         healthy, or nothing is being produced   ← alert on this too
oldest message age           the honest one: how stale is the worst case

Alert on oldest-message-age rather than depth. Depth of 10,000 draining in 30 seconds is fine; depth of 3 stuck for an hour is an incident — Alerting.

Ordering

Global ordering and parallel consumers are mutually exclusive. What you get instead is ordering within a partition key — all messages for order 1234 in sequence, different orders in parallel. Choose the key carefully: too coarse and you’ve serialised everything, too fine and related messages race.

When not to use one

  • The caller needs the result. A queue plus polling for a response is a slow, complicated remote procedure call — REST GraphQL and RPC
  • The work is fast and the failure is the caller’s problem. Validating a postcode doesn’t need a queue
  • You have one server and low volume. A database table with a status column and a cron worker is a queue, is debuggable with SQL, and is often the right amount of machinery
  • It’s hiding a slow dependency you could fix. Queues make latency invisible rather than absent

Where it interacts

  • Webhooks — the receive-then-enqueue pattern is what makes webhook endpoints survivable
  • Race Conditions — queues remove some by serialising per key, and introduce others through out-of-order processing
  • Eventual Consistency — anything behind a queue is stale by definition, so the UI has to be honest about it
  • Observability — propagate the trace context into the message, or every async path becomes an unexplained gap