# Kleppmann and Riccomini update DDIA: storage is moving to object stores, and the trade-offs are moving too

> More and more data systems keep their data in object stores, not on a server's local disk. The old trade-offs are all still there, just in different places, and this book helps you find them.

Bản gốc: https://fdetimes.net/en/books-courses/ddia-second-edition-object-storage-tradeoffs/

Among the changes discussed around the second edition of *Designing Data-Intensive Applications*, one sounds like pure infrastructure: where the data lives. According to an article on HackerNoon, a growing number of systems write to object stores instead of to a file system on a local disk.

For a Forward Deployed Engineer (FDE), that detail matters. It decides whether a customer's cluster can scale quickly, how much data a midnight outage costs, and where the end-of-month bill comes from. That makes the book worth rereading, even if you read the first edition.

## Who wrote it, and what has changed?

Kleppmann's personal website says the book is published by O'Reilly Media, came out in March 2026 and runs to about 670 pages. HackerNoon reports that this time Kleppmann brought in Chris Riccomini, a long-time collaborator, as co-author.

Kleppmann describes the new edition as building on the first, adding new technologies and emerging trends.

According to HackerNoon, the aim was to keep the core intact, weave in new ideas and refresh the detail throughout. In other words, this is a major update, not a new book.

That choice makes sense. From the first edition, the dataintensive.net site presented the book as a map for navigating the diverse and fast-changing world of data storage and processing technologies. A good map does not mark every shop. It shows the main roads.

**Điểm mấu chốt:** Tools change every year. Questions about data trade-offs stay with you for a long time.

## Idea one: every design is a negotiation

The book is organised around five trade-offs that every modern data system has to resolve: scalability, consistency, reliability, efficiency and maintainability. You rarely get all five at once.

FDEs deal with this every day. When a customer says "the dashboard has to be real-time", the real question is how many seconds of stale data they will accept, and what they gain in cost or stability in return. Asking that in customer discovery often saves the team a week of rework.

## Idea two: object stores move the trade-offs

One change the article highlights is that storage increasingly relies on object stores, remote services that are already replicated, rather than local file systems.

To see what follows, picture a three-node database cluster on local disks, with each node holding a copy. At month-end the customer wants to grow to six nodes to run reports. Each new node has to copy the data across before it can do any work, and that copying costs both time and bandwidth.

If the data sits in an object store, the three extra nodes copy nothing. They read directly from the shared store, shut down when they are done, and the provider handles replication. But the trade-off does not disappear. In principle, remote reads are slower than disk reads, and request counts become a line item on the bill.

## How do you bring this into a customer's stack?

The book gives you a vocabulary for trade-offs. To use it on a customer site, you can add a lens the book does not offer: split the stack into three planes. Suppose the customer runs a managed data warehouse with data in object storage.

| Plane | In the warehouse example | Questions to ask the customer |
|---|---|---|
| Control plane | Catalog, table metadata, access permissions, cluster configuration | Who can change the schema? How do you recover if metadata is corrupted? |
| Data plane | Data files in the object store, already replicated | Which region does the data live in? How long is it retained? |
| Compute plane | Query clusters spun up on demand, holding no state | What is the concurrency limit? Who pays for long-running queries? |

Map the book's five trade-offs onto this table and they land in different places. Reliability of the raw data is largely carried by the data plane, while scalability sits in the compute plane. Consistency and maintainability gather in the control plane: when two jobs write to the same table, the catalog decides which one wins, not the disk.

So in the first working session with a customer, ask about the control plane first. It holds the most important state, yet it is the layer people rarely draw.

## Idea three: learn principles first, tools second

The way the book itself was revised, keeping the core and refreshing the detail, doubles as study advice. A year from now, the product names on your CV may be dated. Being able to explain why a system chose consistency over low latency will still be useful.

If you have two to four years of experience and have not read either edition, read the second from the start, slowly, one chapter a week, and compare each chapter with a system you run.

If you have read the first edition, spend your time on the parts where the book adds new technologies and trends.

When reading FDE job descriptions, look for phrases such as "data pipelines" or "distributed systems". Before an interview, prepare a story about a specific trade-off decision you made, and tell it using the book's five concepts.

The book will not help you choose a database for your next project. It helps you ask the right questions before you have to choose, and on a customer site the right question is often worth more than a quick answer.

**Thử ngay tuần này:**

- Take a data system you work on, draw it as three planes (control, data and compute), and note which plane holds state
- For each plane, write one answer: if it dies at 2 a.m., what is lost, and how long does recovery take?
- Add a line to your CV describing a data trade-off decision you made, such as accepting a few seconds of stale reads in exchange for low latency

## Nguồn

- [designing data intensive applications 2e (martin.kleppmann.com)](https://martin.kleppmann.com/2026/03/24/designing-data-intensive-applications-2e.html)

- [Designing Data-Intensive Applications (dataintensive.net)](https://dataintensive.net/)

- [Rethinking Kleppmann's “Designing Data-Intensive Applications”](https://hackernoon.com/lite/rethinking-kleppmanns-designing-data-intensive-applications)
