The Cost Model
ONELINE
Every architecture choice in this book has a price you can compute on a napkin before you build it. Doing that arithmetic early is the highest-leverage habit in the discipline.
This chapter is last because it needs everything before it. It could equally have been first, because the arithmetic here should precede your design, not audit it afterwards.
The good news: the structure is simple and stable. Prices fall constantly, so memorising cents is pointless. Model the structure and plug in today’s numbers.
Two rates, and the one that matters
Pricing is quoted per million tokens, split into input and output — and output costs materially more, commonly three to five times.
That asymmetry is not arbitrary. It is Chapter 4: input is one parallel prefill pass, output is a serial loop with a full forward pass per token.
Which gives you a first instinct worth having: verbose output is expensive in a way verbose input is not. A prompt asking for a terse structured answer is cheaper than one that invites an essay, twice over — fewer output tokens, at the higher rate.
The stack
| Layer | Driven by | Behaviour | Lever |
|---|---|---|---|
| System prompt, tool defs | Fixed scaffolding | Often the majority of input tokens; paid every call | Prompt caching |
| Retrieved context | Chunk count and size | Scales with k; can dwarf everything else |
Chunk budget, k |
| History, scratchpad | Conversation length | Grows unbounded | Compaction, windowing |
| Model tier | Which model | 10–100x spread across tiers | Right-size, route |
| Output length | Verbosity, format | Billed at the higher rate | max_tokens, terse contracts |
| Thinking tokens | Extended reasoning | Output rate, invisible in the response | Gate by task complexity |
| Retries | Errors, guardrail re-runs | Multiplies on failure | Bounded retries, breakers |
| Agent steps | Loop iterations | Multiplies the whole stack per step | Step and budget ceilings |
The last two rows are where surprises live, because they multiply everything above them.
WATCHOUT
Thinking tokens and agent loops are the two worst cost multipliers because they are invisible.
Extended reasoning is billed at the output rate but does not appear in the response — a call returning two sentences can cost an order of magnitude more than it looks. Reported multipliers range from roughly 3x to 15x depending on the task.
Agent loops multiply the entire stack on every step. This is why the only honest unit of measurement is cost per task, not cost per call. A chat turn costs cents; an agentic task can run from tens of cents to several dollars.
The formula
cost_per_request = Σ(layer_input_tokens × input_rate)
+ (output_tokens + thinking_tokens)
× output_rate
cost_per_task = cost_per_request
× expected_steps
× (1 + retry_rate)
with caching applied as a discount on the cacheable share of input.
That is the whole model. Now use it before you build.
NAPKIN — Sizing a support agent before writing any code
Proposal: an agent that answers support questions using retrieval, average four loop steps, 10% retry rate. 50,000 tasks a month.
Per request: 1,500-token system prompt and tools, 3,000 tokens of retrieved chunks, ~800 tokens of accumulated history, 400 tokens out.
Input 5,300 tok @ $3/M = $0.0159
Output 400 tok @ $15/M = $0.0060
Per request ≈ $0.022
Per task x 4 steps x 1.1 ≈ $0.097
Monthly x 50,000 ≈ $4,850
Now apply two levers. Cache the 1,500-token prefix (about 28% of input) at a tenth of the price, and cut retrieval from 3,000 to 1,500 tokens by retrieving 5 chunks instead of 10:
Input 3,650 tok, 1,500 cached ≈ $0.0069
Per request ≈ $0.0129
Per task ≈ $0.057
Monthly ≈ $2,850
41% saved, decided in ten minutes, before a line of code existed.
That is the habit this chapter is selling. Not a spreadsheet — ten minutes of arithmetic that changes the design while changing it is still free.
Almost every runaway bill is one of the nine below.
The nine expensive mistakes
| Anti-pattern | What is happening | Fix |
|---|---|---|
| No prompt caching | Re-paying for the static prefix every call | Stable prefix + provider caching |
| Oversized model | Frontier model on a task a small one handles | Right-size; cascade |
| Unbounded output | No max_tokens, verbose format |
Caps and terse output contracts |
| Reasoning always on | Extended thinking for trivial tasks | Gate by complexity |
| Retry storms | Transient error triggers unbounded retries | Bounded retries, circuit breaker |
| Runaway agent loop | Loop never terminates; errors read as “retry” | Hard ceilings inside the loop |
| Unbounded memory | History accrues without summarisation | Windowing, compaction |
| Long-context stuffing | Giant context as the default retrieval strategy | RAG for large or changing corpora |
| Sync API for offline work | Real-time calls for evals and backfills | Batch endpoints |
Two deserve emphasis.
Right-sizing is the largest single lever available, because the spread across model tiers is 10–100x. Most production traffic is classification, extraction, routing, and simple answers — tasks a small model does well. A cascade runs the cheap model first and escalates only when it is uncertain, which typically routes the large majority of traffic to the cheap path.
NAPKIN — What a cascade is worth
Take a classification-shaped task. On the frontier model it costs about $0.10; on a small model about $0.008 — call it twelve times cheaper. You run 50,000 of them a month.
Frontier for everything: 50,000 x 0.10 = $5,000.
Cascade, with the small model escalating the 20% it is unsure about: every task pays the small model, and one in five pays both.
50,000 x 0.008 = $400, plus 10,000 x 0.10 = $1,000 → $1,400.
72% saved, and the 20% that needed the frontier model still got it.
Now the question that decides whether to build it at all. Cascade costs more than going straight to the frontier model only when 0.008 + (e x 0.10) > 0.10 — that is, when you escalate more than 92% of the time.
At a 12x spread, cheap-first is wrong only if the cheap model is wrong almost always. The expensive part of a cascade was never the extra call; it is the confidence signal that decides when to escalate.
Hard ceilings inside agent loops is the one that prevents catastrophes rather than saving percentages. Monitoring tells you about the money after it is spent. A step limit, a token limit, and a spend limit enforced in the loop stop it.
Attribution, or you cannot manage any of this
If all traffic goes through one untagged API key, you know your total spend and nothing else. You cannot tell which feature, tenant, or customer is expensive, so you cannot make any decision except panic.
Route calls through a gateway that tags every request with feature, tenant, and environment. Then the question stops being “why is the bill high” and becomes “the document-summary feature costs $0.40 per use and is used twice a month” — which is an answerable business question.
What to carry forward
Output costs several times input. Model the layer stack, measure cost per task, and remember that thinking tokens and agent steps multiply invisibly. Cache the stable prefix, right-size the model, cap the output, and put hard ceilings inside every loop. Tag everything.
And do the arithmetic before you build. Ten minutes with this formula will change more designs than any optimisation you apply afterwards.
That closes Part IV, and with it the fifteen mechanisms. One chapter remains, and it is the only one in the book that teaches no mechanism at all. You now have fifteen prices. What you do not have yet is the order in which to spend them.
RECALL
- Why does output cost several times more than input? Tie it to Chapter 4.
- Write the cost-per-task formula from memory.
- Size this before building: 2,000-token prompt, 4,000 tokens retrieved, 600 out, 3 steps, 5% retries, 20,000 tasks a month, $3/M in and $15/M out.
- Why are thinking tokens and agent loops the two most dangerous multipliers?
- Your bill tripled with no traffic change. Name the four likeliest causes.
- Why is per-feature attribution a prerequisite for cost control rather than a reporting nicety?