Grokking AI System Design
The fifteen ideas the rest is built on
How to read this book
This book is fifteen mechanisms and one method.
The claim it makes is that AI system design becomes tractable once a small number of mechanisms are understood well enough that the architecture, the latency, the failure modes, and the bill all fall out of them. Not fifteen topics to be aware of — fifteen things that, once you can compute with them, answer most of the questions you will actually be asked.
A concept earned a chapter here only if it is a mechanism or a trade-off law. Vendor names, model versions, prices, and framework APIs are deliberately absent — real, and wrong within a quarter. They live in the guide this book was distilled from. What is here should still be true after the current model names have been replaced twice.
The chapters chain. They are not independent essays and do not reward skipping. Attention explains why the KV cache exists; the KV cache explains what concurrency costs; chunking sets the ceiling reranking then works under. One system is costed in Chapter 1, made cheaper in Chapter 5, sized in Chapter 15, and designed from its requirements in Chapter 16 — reaching the same number from the other direction.
What is assumed. That you can already design ordinary software: APIs, databases, services, a cache, a queue. Some feel for distributed systems and for probability helps. No machine learning background is needed — tokens, embeddings, and attention are built from nothing in Part I. If you have never shipped a backend system, parts of this will move quickly.
The questions closing each chapter have no printed answers. That is deliberate. Answer them from memory, out loud, badly at first — the only reliable way to find out whether a mechanism has actually landed. Where an answer is not obvious, the guide has it.
Sixteen chapters, five parts. Start at the beginning; the first one is about integers.