Tags: web-dev analytics content

Kleppmann - Designing Data-Intensive Applications - 2017

Author: Martin Kleppmann — 2nd edition with Chris Riccomini
Published: 2017, O’Reilly. 2nd edition February 2026
Format: Book, ~590pp
Read:


The default recommendation for moving past framework knowledge into systems judgement. Known universally as DDIA, and unusual in that it teaches you how to reason about data systems rather than how to operate any of them — which is why a 2017 book was still the standard reference nine years later.


Why it’s the one that gets named

Most systems books either document a specific technology, and expire, or stay so abstract they can’t be acted on. DDIA sits between: it explains the mechanisms underneath the tools, so the knowledge survives the tools.

That’s the reason it became the standard answer to “what should I read to understand backend systems”, and why it appears on reading lists at companies with nothing in common technically. Referred to as DDIA, or “the boar book” after its O’Reilly cover animal.

There is now a second edition (February 2026), co-authored with Chris Riccomini, thoroughly revised — the biggest changes being cloud-native data architectures and the impact of AI workloads. [CHECK: whether the three-part structure below survived the revision, before describing the 2nd edition’s contents specifically.]

Who wrote it

Martin Kleppmann
  built stream-processing infrastructure
  at LinkedIn — co-created Apache Samza
  earlier, co-founded Rapportive

  then a researcher at Cambridge, working
  on CRDTs and local-first software

The industry-to-academia path is what gives the book its character. It cites primary research heavily — most chapters end with dozens of paper references — while staying grounded in what actually breaks in production. Very few books manage both, and it’s the reason the book is trusted by both audiences.

The skew worth knowing: he is a distributed-systems researcher, and the book is most alive in Part II. The material on transactions, replication and consensus is where his expertise is deepest, and it shows.

The organising idea

There is no general-purpose data system. Every tool has bought a set of tradeoffs, and your job is to know which ones you’re buying.

The subtitle is the thesis — the big ideas behind reliable, scalable, and maintainable systems. Those three words are the frame the book opens with and returns to, and they’ve been borrowed widely enough that people use them without knowing the source:

RELIABILITY      it keeps working when
                 things go wrong —
                 hardware, software, humans

SCALABILITY      it keeps working as load
                 grows. Requires defining
                 load and performance first,
                 which is most of the work

MAINTAINABILITY  people can keep working
                 on it — operability,
                 simplicity, evolvability

The move that makes the book unusual is the refusal to recommend. Chapters end without a reference architecture or a “use X” verdict. Reviewers looking for a manual read this as evasion; it’s the deliberate choice, and it’s what makes the content durable — a recommendation would have expired in 2019.

What it put into circulation

The terms and framings people use because of this book, or that it made legible:

TermWhat it means
Reliability · scalability · maintainabilityThe three concerns, as the standard way to open a design discussion
Percentiles over averagesResponse time is a distribution; p95/p99 describe experience, means don’t — Percentiles in Performance
Tail latency amplificationOne slow backend call makes the whole request slow, so tails compound across services
LSM-trees vs B-treesThe write-optimised versus read-optimised storage tradeoff — Indexing
Schema-on-read / schema-on-writeWhether structure is enforced at write time or interpreted at read time
Backward / forward compatibilityOld code reading new data, and new code reading old — Backwards Compatibility
Single-leader · multi-leader · leaderlessThe three replication topologies, and what each costs
Read-your-writes, monotonic reads, consistent prefixThe specific anomalies replication lag produces
Write skew and phantomsThe isolation failures that snapshot isolation doesn’t prevent
Unbundling the databaseDatabases, caches, indexes and stream processors as variations on one idea
Derived data vs system of recordWhat can be rebuilt versus what can’t — the distinction the whole of Part III rests on

The ideas worth being able to explain

Percentiles, and why p99.9

The most immediately applicable thing in the book, and the one that transfers straight into web performance work.

mean response time      hides everything
p50                     the typical request
p95 / p99               the ones people
                        complain about
p99.9                   the ones you'd
                        rather ignore

The argument for caring about extreme tails is the part usually missed: the customers with the slowest requests are frequently the ones with the most data — which is to say the most engaged and most valuable. Optimising the mean optimises for the customers who matter least.

Tail latency amplification is the companion idea: if one page makes twenty backend calls and each has a 1% chance of being slow, most page loads hit at least one slow call. Tails compound; means don’t.

This is the same discipline behind Core Web Vitals being assessed at the 75th percentile rather than the average — the metric was designed by people who had read this argument.

Read-your-writes, which you have already met

Asynchronous replication means a read can hit a replica that hasn’t received your write yet:

t0   user updates basket    → LEADER
t1   page reloads, reads    → REPLICA
                              (lagging)
t2   basket shows the OLD state

The user concludes the site is broken, because from their perspective it is. Read-your-writes consistency is the guarantee that a user always sees their own writes — usually implemented by routing a user’s reads to the leader for a period after they write.

This is a routine ecommerce bug rather than an exotic distributed-systems problem, and it appears the same way in analytics: an event written and immediately queried back may not be there — Eventual Consistency, Tool Discrepancies.

Encoding and evolution

The most underrated chapter, and the one closest to daily work outside backend engineering.

BACKWARD COMPATIBILITY
  new code can read old data
  → needed whenever you deploy

FORWARD COMPATIBILITY
  old code can read new data
  → needed whenever a client is
    out of your control

Every rolling deploy has both running simultaneously. Every mobile app that hasn’t updated is old code reading new data. Every stored analytics event was written by a schema version that no longer exists.

This is the same problem as event schema versioning and API versioning, argued more rigorously than either field usually manages — Event Taxonomy Design, Tracking Plans, Versioning.

The CAP critique

Worth knowing because it’s a genuine, defensible contrarian position and it comes up.

Kleppmann argues CAP is largely of historical interest and its popular formulation is misleading. The “pick two of three” framing implies a menu of choices; in reality it addresses one consistency model (linearizability) and one fault type (network partitions), while saying nothing about latency, other consistency models, or the faults systems actually experience.

The practical version: calling a database “CP” or “AP” tells you almost nothing useful about it — CAP Theorem.

The example everyone cites

Twitter’s home timeline, from Chapter 1. The canonical worked example of scalability reasoning.

APPROACH 1 — compute on read
  a tweet is one insert
  a timeline is a join across everyone
  you follow
  → cheap writes, expensive reads

APPROACH 2 — fan out on write
  each user has a cached timeline
  a tweet is written into every
  follower's cache
  → expensive writes, cheap reads

Twitter moved 1 → 2, then to a HYBRID:
  celebrities with millions of followers
  make fan-out ruinous, so their tweets
  are merged in at read time

The lesson is the one the chapter is actually making: the right architecture depends on the load parameter, and choosing that parameter is the hard part. Here it isn’t tweets per second — it’s the distribution of followers per user, which is where the whole problem lives. Pick the wrong load parameter and you optimise something that isn’t the bottleneck.

That framing generalises well beyond Twitter, which is why it’s the example that gets reused.

How it’s built

First edition, three parts:

PART I    Foundations of Data Systems  (1–4)
          reliability/scalability/maintainability
          data models · storage engines
          encoding and evolution

PART II   Distributed Data  (5–9)
          replication · partitioning
          transactions · what goes wrong
          consistency and consensus

PART III  Derived Data  (10–12)
          batch processing · stream processing
          the future of data systems

Part I is the widely useful part and the one most readers get real value from. Part II is the book’s centre of gravity and its hardest material. Part III is the most speculative — it’s the “unbundling the database” argument, and it’s the section reviewers most often find less useful, partly because it was forecasting a future that has since partly arrived.

The honest verdict

Reception is unusually positive — it’s one of the few technical books with a near-consensus, routinely described as the default recommendation for engineers moving beyond framework knowledge into durable systems judgement.

The recurring criticisms are real but mild:

  • It is dense. Reviewers describe reading it as being as intensive as the systems it covers. It is not a book to get through in a week
  • It refuses to prescribe. No reference architectures, no step-by-step, no “use this” at the end of a chapter. Readers wanting a manual find this evasive — and it’s the deliberate choice that keeps the book relevant
  • The first edition dated on specifics while the concepts held, which is precisely what the second edition exists to fix
  • Part III is the weakest, and the least applicable to most readers

What it’s genuinely good for: the vocabulary, the reasoning patterns, and the ability to ask the right question about an unfamiliar system. The reason it belongs in a CRO vault is narrower but real — percentiles, replication lag, schema evolution and derived data are all things you meet weekly under other names.

Where this vault covers it

Deliberately not repeated here:

The one-paragraph take

“It’s the systems book — the one that teaches you the mechanisms underneath the tools rather than any particular tool, which is why it aged so well. The parts I actually use are Chapter 1 on percentiles, which is the argument behind measuring performance at p75 rather than the mean, and the encoding chapter, which is the clearest treatment of backward and forward compatibility anywhere — it’s the same problem as versioning an event schema. He’s also good on why CAP is a misleading way to describe a database. It’s dense and it deliberately never tells you what to use, which some people find frustrating. There’s a second edition out this year.”