Tags: web-dev concept

CAP Theorem

Date: 2026-08-17


A formal result about distributed systems that is almost always stated wrongly. “Pick two of three” is not what it says, and the practitioners closest to it now argue the framing does more harm than good.


The CAP theorem states that a distributed data store cannot simultaneously guarantee all three of:

CONSISTENCY          every read returns the
                     most recent write
                     (specifically: LINEARIZABILITY)

AVAILABILITY         every request receives
                     a non-error response

PARTITION TOLERANCE  the system continues
                     despite network messages
                     being lost between nodes

Why “pick two” is wrong

Partition tolerance is not optional. Networks fail — cables, switches, routing, cloud zones. A distributed system cannot choose not to experience partitions; it can only choose how it behaves during one.

So the real statement is much narrower:

NOT:  pick 2 of 3

BUT:  WHEN A PARTITION OCCURS,
      choose consistency or availability

      no partition → you can have both

Most of the time there is no partition, and the trade-off doesn’t apply at all. A system labelled “AP” is not perpetually inconsistent; it’s a system that will prefer to answer rather than refuse during a partition.

The choice, concretely

A NODE IS CUT OFF FROM ITS PEERS.
A request arrives. It cannot verify it
has the latest data.

CP — choose consistency
  refuse to answer
  → correct, unavailable
  → banking, stock, bookings

AP — choose availability
  answer with what it has
  → available, possibly stale
  → product pages, feeds, catalogues

Neither is right in general. Serving a stale product description is fine; selling stock you don’t have is not.

Why it’s now considered a poor framing

Kleppmann and others have argued the theorem is largely of historical interest, and the criticism is specific and fair — Kleppmann - Designing Data-Intensive Applications - 2017:

ONE consistency model
  CAP's "C" is linearizability only.
  There is a whole spectrum of weaker
  models it says nothing about

ONE fault type
  network partitions only. Not slow
  nodes, crashed nodes, clock skew,
  or the far more common case of a
  node that is merely SLOW

NO latency
  a system that answers correctly after
  ten seconds is "available" by CAP and
  useless in practice

BINARY labels
  calling a database "CP" or "AP" tells
  you almost nothing about how it
  behaves

“Available” in CAP means “eventually returns something”, with no time bound. That’s not what anyone means by available.

PACELC, which is the more useful version

An extension that adds the case CAP ignores — the normal one:

IF Partition:
    choose Availability or Consistency
ELSE (no partition):
    choose Latency or Consistency

The “else” branch is where systems spend 99.9% of their time. Even with a healthy network, strong consistency costs coordination, and coordination costs latency. That trade is continuous and always present, which makes it more useful to reason about than a failure mode you see twice a year — Eventual Consistency.

What to actually take from it

  • Consistency is per-data, not per-system. Stock strongly consistent, product copy eventually consistent, in the same application
  • Ask what happens during a partition when evaluating a datastore — the answer is more informative than a two-letter label
  • Latency is the everyday version of the trade. Every cache and every read replica is choosing speed over freshness
  • If someone says “we chose AP”, the useful follow-up is which data, and what the user sees when it’s stale

Know it well enough to recognise the misstatement, and to talk about consistency in terms of specific guarantees for specific data rather than in three letters.