Putting It Together
ONELINE
Design is not choosing components. It is asking the cheapest disqualifying question first, so that every later choice is made against a constraint whose price you already know.
Fifteen chapters, fifteen mechanisms. Not one of them tells you what to build.
What they give you is a price list. Attention tells you what a long prompt costs. The KV cache tells you how many users fit on a box. The agent loop tells you what a twentieth step does to your success rate. Every chapter converts a design choice into a number you can compute before you write any code.
This chapter spends them, in order, on one system.
The system
An enterprise support assistant. It answers questions from private documentation, looks up the current state of a customer’s account, opens tickets, escalates what it cannot resolve, and cites its sources. Several tenants share one deployment. Three-second p95.
You have met its bill twice. Chapter 1 costed its simplest ancestor at $1,200 a month. Chapter 5 cut that to $660 with prompt caching. Chapter 15 sized the full agent at $4,850, then at $2,850. This chapter is the reasoning that produces the thing Chapter 15 priced.
Eight questions. The order matters more than any single answer, because each one narrows what the next is allowed to be.
1. Does this need a model at all?
Ask first, because it is free to ask and expensive to skip.
Sort the request surface before designing for it. “Reset my password” is a routing problem — a classifier and a link. “What is my balance” is a database query with a sentence wrapped round it. “Why was I charged twice in March, and does your refund policy cover it” is the one that genuinely needs a model: it requires reading a policy, reading an account, and reasoning about both together.
Chapter 10 put this on a ladder from L0 to L4, with the warning that most problems solved at L2 should have been L0 or L1. That applies to the product, not only to agents. A fixed path is cheaper, faster, testable, and debuggable. You trade all of it for flexibility, and you should only pay when the path truly cannot be known in advance.
Perhaps a third of this system’s traffic never needs a model. Routing it away is the largest single cost reduction available here, and it happens before the other seven questions have been asked.
WATCHOUT
The temptation runs the other way, because one prompt handles all three cases and the demo works. That is exactly how teams ship the general path for everything and meet the bill and the latency later.
Classifying first costs one small-model call. It buys a deterministic path for the traffic that deserves one, and a shorter, sharper prompt for the traffic that does not.
2. Name what you are optimising, and what you will trade
Latency, throughput, availability, scalability, security, maintainability — all still apply, unchanged. AI systems add attributes conventional architecture has no vocabulary for, and each already has a chapter behind it.
| Attribute | The question it asks | Chapter |
|---|---|---|
| Faithfulness | Is the answer supported by what we retrieved? | 13 |
| Context recall | Did the right material reach the model at all? | 7, 13 |
| Refusal quality | Does it decline when it should, and say so? | 14 |
| Task success rate | Did the whole task finish, not just one step? | 10 |
| Cost per task | What did it cost, including steps and retries? | 15 |
| Reproducibility | Does the same input behave the same way twice? | 4 |
The list is not the point. Ranking it is, because several of these oppose each other directly. Faithfulness trades against helpfulness: a system that refuses whenever it is less than certain is perfectly faithful and largely useless. Cost per task trades against success rate: more steps, more retries, and a bigger model all buy accuracy.
For a support assistant the ranking writes itself. Faithfulness first — a confidently invented refund policy is worse than no answer at all. Refusal quality second. Cost per task third. Raw helpfulness last, because escalating to a human is a perfectly good outcome and the cheapest safety net available.
Write that order down. Most of what follows is resolved by it.
3. Private knowledge: retrieval, or just context?
Chapter 6 gave the regime rule, and it deserves an honest answer before anyone builds a pipeline. Under roughly 50,000 tokens of stable content, put it in a cached prompt: no index, no chunking decision, no embedding migration, and Chapter 5’s caching makes it cheap.
This corpus is gigabytes and changes daily, so we are in the other regime. But notice what the question did. It made build RAG a conclusion rather than an assumption, and it named the condition under which the conclusion would flip.
4. Which retrieval shape, and what does it cost in latency?
Chapter 7’s ladder, climbed slowly. Naive RAG is a demo. Advanced RAG — hybrid search and reranking — is the sensible production default. Agentic and graph retrieval are answers to measured failures you have not had yet.
Advanced, then. Which settles two decisions already made for you.
Chunking (Chapter 8) is the one nothing downstream can undo. This corpus is prose with tables in it, so: recursive splitting into roughly 300-token children grouped under 1,500-token parents, index the children and return the parents, never split a table from its header, and contextual prefixes — because a chunk reading “this does not apply to enterprise customers” is worse than useless alone.
Two-stage retrieval (Chapter 9) closes gaps 1 and 2 together. Hybrid dense and sparse is not optional here: error codes and product identifiers are precisely what dense retrieval cannot see. Fuse by RRF, then rerank the top 100 with a cross-encoder.
Now spend the latency budget, because three seconds is a constraint rather than a hope.
NAPKIN — Three-second p95, of what exactly?
Time to first token, adding up the stages:
input guardrail 80 ms
classify and route 150 ms
hybrid retrieval 120 ms dense + sparse
rerank 100 400 ms ~4 ms each (Ch 9)
prefill ~5,300 tok 300 ms
---------
TTFT ~1,050 ms
Comfortable. Now the complete response: 400 output tokens at roughly 60 tokens per second is another 6.7 seconds. Total ≈ 7.8 s.
So the SLO means one of two entirely different things. If p95 is measured to first token, you have two seconds of headroom and the architecture above is fine. If it is measured to the last token, you are 2.5x over, and the only levers are fewer output tokens or a faster model — Chapter 4, decode is serial and nothing here changes that.
Settle which number the SLO means before you design, not after. One reading you can meet with the design you already have; the other rewrites it.
It also rules something out: an agent loop streams nothing to the user until it stops looping. Whatever runs on the interactive path cannot be a loop.
5. Workflow or agent?
Chapter 10’s arithmetic decides this and it is unforgiving. At 95% per step, four steps succeed 81% of the time and twenty steps succeed 36%.
So: as little agent as the task will admit.
Answering a documentation question is a pipeline, not a loop — retrieve, rerank, generate, check the citations. Looking up an account is one tool call. Creating a ticket is a fixed sequence: validate, create, confirm. All L0 and L1, all testable, all debuggable.
The loop earns its place in exactly one branch: the question that needs several lookups whose shape depends on what the earlier ones returned. Bound it at four steps — what Chapter 15 priced, and what the arithmetic tolerates — and make hitting the cap an escalation rather than an apology.
6. Where are the boundaries and the gates?
Three questions, three different answers, all from Part III and Chapter 14.
What the system can do, and who decides. Tools sit behind the capability boundary of Chapter 11. The tenant comes from the authenticated session and never from an argument the model filled in. The refund tool has a ceiling, and the ceiling lives in the tool rather than in the prompt, because a prompt instruction is a suggestion and a credential is a guarantee.
What needs a human. Chapter 10’s test is reversibility, not importance. Reading documentation, reading an account, searching: let them run. Issuing a refund, closing an account, mailing a customer: gate them. Put the gate on the tool so every path that ever reaches it inherits the gate.
What happens when the content is hostile. Support tickets are attacker-writable text and this system reads them. Chapter 14 is blunt that no prompt fixes this. Separate the trust domains: whatever reads ticket content does not hold the credential that can send mail. Guardrail the input, the actions, and the output, because each catches something the others do not.
7. How will you know it works?
Chapter 13: the eval is the product, and it has to exist before the system does.
Score retrieval and generation separately — they fail independently and look identical from outside. Context recall and precision say whether the right chunks arrived; faithfulness and answer relevancy say what the model did with them. High recall with low faithfulness is a generation problem. The reverse is a retrieval problem. One number for “quality” tells you neither.
Put correctly refused and correctly escalated in the eval set as passes. Step 2 ranked refusal quality second, and a metric that punishes declining will quietly train the system out of the behaviour you asked for.
Then trace everything, sampled, retaining every failure. When faithfulness drops next quarter, the trace is the difference between ten minutes and a week.
8. What does it cost?
Last, because only now is there something to cost.
Chapter 15 did this arithmetic in full: a 1,500-token cached prefix, 3,000 tokens of retrieved chunks, 800 tokens of history, 400 out, four steps, 10% retries, 50,000 tasks a month. $4,850, falling to $2,850 once the prefix is cached and retrieval is cut from ten chunks to five.
Two things are worth noticing about where that number came from.
It is the same number, reached from the opposite direction. Chapter 15 started from a proposal and priced it. This chapter started from a requirement and derived it. That they meet is the point of the whole book: the mechanisms are constraints, and constraints do not care which end you enter from.
And the largest saving is not in the list. Question 1 routed roughly a third of traffic away from the model entirely, before any of this arithmetic applied. The cheapest token is still the one you never send.
The system those eight questions produced
Two things in that picture are worth more than the boxes.
The line that goes around the model. Question 1 sent roughly a third of traffic down the deterministic path, where it never becomes a token. That branch is the largest cost decision in the diagram, and it was settled before any of the other seven questions were asked.
The trace underneath. It is not a stage. It is the record every stage writes to, and Chapter 13’s evals, Chapter 15’s cost attribution, and Chapter 14’s incident review all read the same rows. Systems that add it last never quite get it, because you cannot trace requests that have already been served.
What is absent is just as decided. There is no agent loop on the interactive path — the latency napkin ruled it out, because a loop streams nothing until it stops. There is no memory tier beyond the session, because nothing in the requirements asked the assistant to remember a customer between conversations, and Chapter 12 is blunt that a durable store you did not need is a privacy surface you now own.
The sequence, and the trades inside it
The shapes, and where each one is taught
Eight shapes cover most of what you will build. None is new — seeing them together is how you notice which one you are reaching for, and what it charges.
| Shape | The right answer when | What it costs you | Ch. |
|---|---|---|---|
| Single call | One question, no private data, no action | Nothing — start here | 1-5 |
| Classify and route | The request surface is mixed | A classifier, and a taxonomy to keep current | 16 |
| RAG pipeline | Knowledge is private, large, or changing | An index, and a chunking decision you cannot undo | 7-9 |
| Tool-using agent | The path depends on what earlier steps return | p^n reliability and context that grows every turn |
10-11 |
| Capability layer | A tool outlives the host that first used it | One more service, and a protocol to track | 11 |
| Approval gate | The action cannot be taken back | Throughput, and a human in the path | 10, 14 |
| Model cascade | A small model handles most of the traffic | A confidence signal you have to build and trust | 15 |
| Sub-agent isolation | One step would flood the parent’s context | Coordination, and a summary that drops detail | 6 |
Most production systems are two or three of these composed rather than one. The support assistant is classify-and-route wrapping a RAG pipeline, with a bounded agent on a single branch and an approval gate on a single tool.
Every step above is a trade this book has already priced. Collected in one place, so you can argue about them before you build rather than after:
| Decision | Cheap side | Expensive side | What you are trading | Ch. |
|---|---|---|---|---|
| Knowledge | Cached context | Retrieval pipeline | Index upkeep vs corpus size and freshness | 6, 7 |
| Retrieval | Dense only | Hybrid + rerank | 100-400 ms and a second index vs gaps 1 and 2 | 9 |
| Chunk size | Large chunks | Small, parent-child | Precision at search vs context at generation | 8 |
| Control flow | Workflow | Agent loop | Predictability and p^n vs paths you cannot enumerate |
10 |
| Model | Small, cascade | Frontier everywhere | 10-100x cost vs quality on the hard tail | 15 |
| Actions | Autonomous | Human approval | Throughput vs irreversibility | 10, 14 |
| Output | Free text | Structured | Expressiveness vs something you can validate | 13, 14 |
| Memory | Context only | Durable store | Simplicity vs continuity, and a privacy surface | 12 |
| Evaluation | Offline only | Offline and online | Iteration speed vs knowing what production does | 13 |
| Failure | Retry | Degrade | Latency and cost vs a deliberately worse answer | 14 |
There is no column that is always right. There is only the ranking you wrote down in question 2, and what it says about each row.
What to carry forward
The sequence is the deliverable, not the architecture it produced. Ours suits a support assistant with those requirements and that ranking; change the ranking and the same eight questions produce a different system.
Ask them cheapest first. Most designs are decided by questions 1 and 2, and teams reach them last — after the framework is chosen, the pipeline is built, and the constraint that would have ruled it out has become expensive to discover.
That is the whole method. The fifteen mechanisms give you the prices. This chapter gives you the order.
The end, and the beginning
Fifteen ideas: tokens, embeddings, attention, the generation loop, the KV cache, the context budget, the retrieval gap, chunking, two-stage retrieval, the agent loop, tools, memory, evaluation, failure, and cost. Then one method for spending them.
They are the load-bearing ones — the concepts that will still be true when the model names on today’s benchmark have been replaced twice over. Everything else in this field is a specific, dated instance of one of them.
The guide this book was distilled from has the specifics: the vendors, the frameworks, the current numbers, the case studies. Go there when you need today’s answer. Come back here when you need to know why it is the answer.
RECALL
- Why does “does this need a model at all” come before every other question?
- Rank the AI quality attributes for a medical triage assistant. Where does your ranking differ from the support assistant’s, and what changes downstream?
- A three-second p95 target is given with no further detail. What do you ask, and how does each answer change the design?
- Your corpus is a 40-page handbook that changes twice a year. Work question 3 and justify the answer.
- Which parts of the support assistant are L0 or L1, and which single branch justifies a loop?
- Give the reversibility test for gating a tool, and apply it to: search tickets, refund an order, read an account, close an account.
- Your eval set scores a correct refusal as a failure. Name two things that will happen over the next quarter.