Tags: web-dev concept

Event-Driven Architecture

Date: 2026-08-17


Services announce facts about what happened rather than instructing each other what to do next. It removes the sender’s knowledge of who’s listening, which is the coupling that makes systems hard to change — and replaces it with nobody knowing what the system does end to end.


Event-driven architecture is a design where components communicate by publishing events — records that something happened — which other components subscribe to, rather than by calling each other directly.

Commands versus events

The whole distinction is tense.

COMMAND                              EVENT

"send the confirmation email"        "order 1234 was placed"

imperative, present tense            declarative, past tense
one named recipient                  no named recipient
sender knows what should happen      sender knows what did happen
sender knows if it worked            sender never finds out

A command has one handler; an event has zero or more. That’s the difference that matters at runtime, and the reason adding a consumer is free.

What it removes

ORCHESTRATED                         EVENT-DRIVEN

checkout service:                    checkout service:
  call warehouse                       publish order.placed
  call email service                        │
  call CRM                                  ├──→ warehouse   (reserves stock)
  call analytics                            ├──→ email       (sends receipt)
  call loyalty                              ├──→ CRM         (updates record)
                                            ├──→ analytics   (records revenue)
adding "loyalty points" means               └──→ loyalty     (awards points)
editing and redeploying checkout
                                     adding a consumer touches nothing upstream
checkout is down if any callee is    checkout doesn't know they exist

Notice what the left-hand column makes checkout responsible for: the business process. Every new step in “what happens after an order” is a change to the checkout service, which is why that file ends up two thousand lines long and owned by six teams.

What it costs

  • No one place describes the flow. “What happens when an order is placed” is answerable only by searching every service for a subscription. This is the honest, permanent cost, and documentation does not fix it — Observability with distributed tracing partly does
  • Debugging is archaeology. A failure three consumers deep surfaces as a missing side effect an hour later, with no stack trace connecting it to the cause
  • The event schema is a public API. Every consumer depends on the shape, you don’t know who they all are, and you can’t change a field without breaking someone — Backwards Compatibility, Deprecation
  • Everything is eventually consistent. The order exists and the loyalty points don’t, for a while. The UI has to admit that rather than pretend
  • Testing is integration testing. Unit tests prove a consumer handles a message; nothing proves the choreography is correct

Designing the events

Name them as facts that already happened: order.placed, payment.captured, stock.reserved. A name like order.needs_email is a command wearing an event’s clothes, and it re-couples the publisher to a consumer.

Thin or fat is a real decision:

Thin (ID only)Fat (full state)
Payload{ order_id: "1234" }The whole order
Consumer mustCall back for detailsNothing
CouplingTo your APITo your schema
LoadN callbacks per eventNone
StalenessAlways currentSnapshot at publish time
ReplayGets today’s data — often wrongGets what was true then — usually right

Fat events with an explicit version field are the better default for anything crossing a team boundary, because the callback storm from thin events lands on the publisher at exactly the moment it’s busiest.

Every event needs: a unique ID (deduplication), a timestamp (ordering), a version (schema evolution), and a correlation ID (tracing one customer action across everything it triggered).

Transactions across services

The hard part. There’s no distributed transaction, so a multi-step process needs a saga: a sequence of local transactions where each step publishes an event triggering the next, and every step has a compensating action for rolling back.

order.placed → reserve stock → capture payment → arrange shipping
                    │                 │
                    │                 └─ fails → release stock  (compensation)
                    └─ fails → cancel order, refund nothing yet

Compensation is not rollback. You can’t un-send an email or un-charge a card without a visible refund — the customer sees the failed path. Design the order of steps so the irreversible ones happen last.

Where it’s the wrong shape

  • The caller needs an answer. “Is this discount code valid” is a question, not a fact — REST GraphQL and RPC
  • A monolith. In-process function calls already give you ordering, transactions and a stack trace, all of which you’d be giving up for nothing — Monolith vs Services
  • Fewer than three consumers of the same event. Below that, the indirection costs more than the coupling it removes

Where it interacts

  • Message Queues — the transport underneath; the log shape is what makes replay and late-joining consumers possible
  • Webhooks — the same idea across an organisational boundary
  • Experiment Assignment Tracking and Event Taxonomy Design — analytics events and architectural events are different things with the same name, and conflating them produces a stream that serves neither