Tags: web-dev analytics content
Kleppmann - Designing Data-Intensive Applications - 2017
Author: Martin Kleppmann — 2nd edition with Chris Riccomini
Published: 2017, O’Reilly. 2nd edition February 2026
Format: Book, ~590pp
Read:
The default recommendation for moving past framework knowledge into systems judgement. Known universally as DDIA, and unusual in that it teaches you how to reason about data systems rather than how to operate any of them — which is why a 2017 book was still the standard reference nine years later.

Why it’s the one that gets named
Most systems books either document a specific technology, and expire, or stay so abstract they can’t be acted on. DDIA sits between: it explains the mechanisms underneath the tools, so the knowledge survives the tools.
That’s the reason it became the standard answer to “what should I read to understand backend systems”, and why it appears on reading lists at companies with nothing in common technically. Referred to as DDIA, or “the boar book” after its O’Reilly cover animal.
There is now a second edition (February 2026), co-authored with Chris Riccomini, thoroughly revised — the biggest changes being cloud-native data architectures and the impact of AI workloads. [CHECK: whether the three-part structure below survived the revision, before describing the 2nd edition’s contents specifically.]
Who wrote it
Martin Kleppmann
built stream-processing infrastructure
at LinkedIn — co-created Apache Samza
earlier, co-founded Rapportive
then a researcher at Cambridge, working
on CRDTs and local-first software
The industry-to-academia path is what gives the book its character. It cites primary research heavily — most chapters end with dozens of paper references — while staying grounded in what actually breaks in production. Very few books manage both, and it’s the reason the book is trusted by both audiences.
The skew worth knowing: he is a distributed-systems researcher, and the book is most alive in Part II. The material on transactions, replication and consensus is where his expertise is deepest, and it shows.
The organising idea
There is no general-purpose data system. Every tool has bought a set of tradeoffs, and your job is to know which ones you’re buying.
The subtitle is the thesis — the big ideas behind reliable, scalable, and maintainable systems. Those three words are the frame the book opens with and returns to, and they’ve been borrowed widely enough that people use them without knowing the source:
RELIABILITY it keeps working when
things go wrong —
hardware, software, humans
SCALABILITY it keeps working as load
grows. Requires defining
load and performance first,
which is most of the work
MAINTAINABILITY people can keep working
on it — operability,
simplicity, evolvability
The move that makes the book unusual is the refusal to recommend. Chapters end without a reference architecture or a “use X” verdict. Reviewers looking for a manual read this as evasion; it’s the deliberate choice, and it’s what makes the content durable — a recommendation would have expired in 2019.
What it put into circulation
The terms and framings people use because of this book, or that it made legible:
| Term | What it means |
|---|---|
| Reliability · scalability · maintainability | The three concerns, as the standard way to open a design discussion |
| Percentiles over averages | Response time is a distribution; p95/p99 describe experience, means don’t — Percentiles in Performance |
| Tail latency amplification | One slow backend call makes the whole request slow, so tails compound across services |
| LSM-trees vs B-trees | The write-optimised versus read-optimised storage tradeoff — Indexing |
| Schema-on-read / schema-on-write | Whether structure is enforced at write time or interpreted at read time |
| Backward / forward compatibility | Old code reading new data, and new code reading old — Backwards Compatibility |
| Single-leader · multi-leader · leaderless | The three replication topologies, and what each costs |
| Read-your-writes, monotonic reads, consistent prefix | The specific anomalies replication lag produces |
| Write skew and phantoms | The isolation failures that snapshot isolation doesn’t prevent |
| Unbundling the database | Databases, caches, indexes and stream processors as variations on one idea |
| Derived data vs system of record | What can be rebuilt versus what can’t — the distinction the whole of Part III rests on |
The ideas worth being able to explain
Percentiles, and why p99.9
The most immediately applicable thing in the book, and the one that transfers straight into web performance work.
mean response time hides everything
p50 the typical request
p95 / p99 the ones people
complain about
p99.9 the ones you'd
rather ignore
The argument for caring about extreme tails is the part usually missed: the customers with the slowest requests are frequently the ones with the most data — which is to say the most engaged and most valuable. Optimising the mean optimises for the customers who matter least.
Tail latency amplification is the companion idea: if one page makes twenty backend calls and each has a 1% chance of being slow, most page loads hit at least one slow call. Tails compound; means don’t.
This is the same discipline behind Core Web Vitals being assessed at the 75th percentile rather than the average — the metric was designed by people who had read this argument.
Read-your-writes, which you have already met
Asynchronous replication means a read can hit a replica that hasn’t received your write yet:
t0 user updates basket → LEADER
t1 page reloads, reads → REPLICA
(lagging)
t2 basket shows the OLD state
The user concludes the site is broken, because from their perspective it is. Read-your-writes consistency is the guarantee that a user always sees their own writes — usually implemented by routing a user’s reads to the leader for a period after they write.
This is a routine ecommerce bug rather than an exotic distributed-systems problem, and it appears the same way in analytics: an event written and immediately queried back may not be there — Eventual Consistency, Tool Discrepancies.
Encoding and evolution
The most underrated chapter, and the one closest to daily work outside backend engineering.
BACKWARD COMPATIBILITY
new code can read old data
→ needed whenever you deploy
FORWARD COMPATIBILITY
old code can read new data
→ needed whenever a client is
out of your control
Every rolling deploy has both running simultaneously. Every mobile app that hasn’t updated is old code reading new data. Every stored analytics event was written by a schema version that no longer exists.
This is the same problem as event schema versioning and API versioning, argued more rigorously than either field usually manages — Event Taxonomy Design, Tracking Plans, Versioning.
The CAP critique
Worth knowing because it’s a genuine, defensible contrarian position and it comes up.
Kleppmann argues CAP is largely of historical interest and its popular formulation is misleading. The “pick two of three” framing implies a menu of choices; in reality it addresses one consistency model (linearizability) and one fault type (network partitions), while saying nothing about latency, other consistency models, or the faults systems actually experience.
The practical version: calling a database “CP” or “AP” tells you almost nothing useful about it — CAP Theorem.
The example everyone cites
Twitter’s home timeline, from Chapter 1. The canonical worked example of scalability reasoning.
APPROACH 1 — compute on read
a tweet is one insert
a timeline is a join across everyone
you follow
→ cheap writes, expensive reads
APPROACH 2 — fan out on write
each user has a cached timeline
a tweet is written into every
follower's cache
→ expensive writes, cheap reads
Twitter moved 1 → 2, then to a HYBRID:
celebrities with millions of followers
make fan-out ruinous, so their tweets
are merged in at read time
The lesson is the one the chapter is actually making: the right architecture depends on the load parameter, and choosing that parameter is the hard part. Here it isn’t tweets per second — it’s the distribution of followers per user, which is where the whole problem lives. Pick the wrong load parameter and you optimise something that isn’t the bottleneck.
That framing generalises well beyond Twitter, which is why it’s the example that gets reused.
How it’s built
First edition, three parts:
PART I Foundations of Data Systems (1–4)
reliability/scalability/maintainability
data models · storage engines
encoding and evolution
PART II Distributed Data (5–9)
replication · partitioning
transactions · what goes wrong
consistency and consensus
PART III Derived Data (10–12)
batch processing · stream processing
the future of data systems
Part I is the widely useful part and the one most readers get real value from. Part II is the book’s centre of gravity and its hardest material. Part III is the most speculative — it’s the “unbundling the database” argument, and it’s the section reviewers most often find less useful, partly because it was forecasting a future that has since partly arrived.
The honest verdict
Reception is unusually positive — it’s one of the few technical books with a near-consensus, routinely described as the default recommendation for engineers moving beyond framework knowledge into durable systems judgement.
The recurring criticisms are real but mild:
- It is dense. Reviewers describe reading it as being as intensive as the systems it covers. It is not a book to get through in a week
- It refuses to prescribe. No reference architectures, no step-by-step, no “use this” at the end of a chapter. Readers wanting a manual find this evasive — and it’s the deliberate choice that keeps the book relevant
- The first edition dated on specifics while the concepts held, which is precisely what the second edition exists to fix
- Part III is the weakest, and the least applicable to most readers
What it’s genuinely good for: the vocabulary, the reasoning patterns, and the ability to ask the right question about an unfamiliar system. The reason it belongs in a CRO vault is narrower but real — percentiles, replication lag, schema evolution and derived data are all things you meet weekly under other names.
Where this vault covers it
Deliberately not repeated here:
- Percentiles in Performance · Core Web Vitals · Latency and Bandwidth
- Eventual Consistency · CAP Theorem · Transactions and ACID
- Indexing · Query Planning · SQL vs NoSQL
- Backwards Compatibility · Versioning · Serialisation Formats
- Warehouse-First Analytics · Event Taxonomy Design
The one-paragraph take
“It’s the systems book — the one that teaches you the mechanisms underneath the tools rather than any particular tool, which is why it aged so well. The parts I actually use are Chapter 1 on percentiles, which is the argument behind measuring performance at p75 rather than the mean, and the encoding chapter, which is the clearest treatment of backward and forward compatibility anywhere — it’s the same problem as versioning an event schema. He’s also good on why CAP is a misleading way to describe a database. It’s dense and it deliberately never tells you what to use, which some people find frustrating. There’s a second edition out this year.”