The Cost Model

Lesson 15 / 16 · updated 2026-10-01 · 9 min


ONELINE

Every architecture choice in this book has a price you can compute on a napkin before you build it. Doing that arithmetic early is the highest-leverage habit in the discipline.

METAPHOR

A menu where every dish lists its price per serving.

This chapter is last because it needs everything before it. It could equally have been first, because the arithmetic here should precede your design, not audit it afterwards.

The good news: the structure is simple and stable. Prices fall constantly, so memorising cents is pointless. Model the structure and plug in today’s numbers.

Two rates, and the one that matters

Pricing is quoted per million tokens, split into input and output — and output costs materially more, commonly three to five times.

That asymmetry is not arbitrary. It is Chapter 4: input is one parallel prefill pass, output is a serial loop with a full forward pass per token.

Which gives you a first instinct worth having: verbose output is expensive in a way verbose input is not. A prompt asking for a terse structured answer is cheaper than one that invites an essay, twice over — fewer output tokens, at the higher rate.

The stack

Cost stackFlowchart with 6 labelled stages: System prompt and tool defs - CACHE IT; Retrieved context - scales with k; History and scratchpad - COMPACT IT; Output and thinking tokens - 3-5x input; = cost per request, x steps, x retries; = COST PER TASK. Connections: System prompt and tool defs - CACHE IT leads to Retrieved context - scales with k; Retrieved context - scales with k leads to History and scratchpad - COMPACT IT; History and scratchpad - COMPACT IT leads to Output and thinking tokens - 3-5x input; Output and thinking tokens - 3-5x input leads to = cost per request, x steps, x retries; = cost per request, x steps, x retries leads to = COST PER TASK.System prompt and tool defs - CACHEITRetrieved context - scales with kHistory and scratchpad - COMPACT ITOutput and thinking tokens - 3-5x input= cost per request, x steps, x retries= COST PER TASK
Every request is a stack of layers. Each has its own driver and its own lever.
Layer Driven by Behaviour Lever
System prompt, tool defs Fixed scaffolding Often the majority of input tokens; paid every call Prompt caching
Retrieved context Chunk count and size Scales with k; can dwarf everything else Chunk budget, k
History, scratchpad Conversation length Grows unbounded Compaction, windowing
Model tier Which model 10–100x spread across tiers Right-size, route
Output length Verbosity, format Billed at the higher rate max_tokens, terse contracts
Thinking tokens Extended reasoning Output rate, invisible in the response Gate by task complexity
Retries Errors, guardrail re-runs Multiplies on failure Bounded retries, breakers
Agent steps Loop iterations Multiplies the whole stack per step Step and budget ceilings

The last two rows are where surprises live, because they multiply everything above them.

WATCHOUT

Thinking tokens and agent loops are the two worst cost multipliers because they are invisible.

Extended reasoning is billed at the output rate but does not appear in the response — a call returning two sentences can cost an order of magnitude more than it looks. Reported multipliers range from roughly 3x to 15x depending on the task.

Agent loops multiply the entire stack on every step. This is why the only honest unit of measurement is cost per task, not cost per call. A chat turn costs cents; an agentic task can run from tens of cents to several dollars.

The formula

cost_per_request = Σ(layer_input_tokens × input_rate)
                 + (output_tokens + thinking_tokens)
                   × output_rate

cost_per_task    = cost_per_request
                 × expected_steps
                 × (1 + retry_rate)

with caching applied as a discount on the cacheable share of input.

That is the whole model. Now use it before you build.

NAPKIN — Sizing a support agent before writing any code

Proposal: an agent that answers support questions using retrieval, average four loop steps, 10% retry rate. 50,000 tasks a month.

Per request: 1,500-token system prompt and tools, 3,000 tokens of retrieved chunks, ~800 tokens of accumulated history, 400 tokens out.

Input    5,300 tok @ $3/M    =  $0.0159
Output     400 tok @ $15/M   =  $0.0060
Per request                  ≈  $0.022
Per task    x 4 steps x 1.1  ≈  $0.097
Monthly     x 50,000         ≈  $4,850

Now apply two levers. Cache the 1,500-token prefix (about 28% of input) at a tenth of the price, and cut retrieval from 3,000 to 1,500 tokens by retrieving 5 chunks instead of 10:

Input    3,650 tok, 1,500 cached  ≈  $0.0069
Per request                       ≈  $0.0129
Per task                          ≈  $0.057
Monthly                           ≈  $2,850

41% saved, decided in ten minutes, before a line of code existed.

That is the habit this chapter is selling. Not a spreadsheet — ten minutes of arithmetic that changes the design while changing it is still free.

Almost every runaway bill is one of the nine below.

The nine expensive mistakes

Anti-pattern What is happening Fix
No prompt caching Re-paying for the static prefix every call Stable prefix + provider caching
Oversized model Frontier model on a task a small one handles Right-size; cascade
Unbounded output No max_tokens, verbose format Caps and terse output contracts
Reasoning always on Extended thinking for trivial tasks Gate by complexity
Retry storms Transient error triggers unbounded retries Bounded retries, circuit breaker
Runaway agent loop Loop never terminates; errors read as “retry” Hard ceilings inside the loop
Unbounded memory History accrues without summarisation Windowing, compaction
Long-context stuffing Giant context as the default retrieval strategy RAG for large or changing corpora
Sync API for offline work Real-time calls for evals and backfills Batch endpoints

Two deserve emphasis.

Right-sizing is the largest single lever available, because the spread across model tiers is 10–100x. Most production traffic is classification, extraction, routing, and simple answers — tasks a small model does well. A cascade runs the cheap model first and escalates only when it is uncertain, which typically routes the large majority of traffic to the cheap path.

NAPKIN — What a cascade is worth

Take a classification-shaped task. On the frontier model it costs about $0.10; on a small model about $0.008 — call it twelve times cheaper. You run 50,000 of them a month.

Frontier for everything: 50,000 x 0.10 = $5,000.

Cascade, with the small model escalating the 20% it is unsure about: every task pays the small model, and one in five pays both.

50,000 x 0.008 = $400, plus 10,000 x 0.10 = $1,000 → $1,400.

72% saved, and the 20% that needed the frontier model still got it.

Now the question that decides whether to build it at all. Cascade costs more than going straight to the frontier model only when 0.008 + (e x 0.10) > 0.10 — that is, when you escalate more than 92% of the time.

At a 12x spread, cheap-first is wrong only if the cheap model is wrong almost always. The expensive part of a cascade was never the extra call; it is the confidence signal that decides when to escalate.

Hard ceilings inside agent loops is the one that prevents catastrophes rather than saving percentages. Monitoring tells you about the money after it is spent. A step limit, a token limit, and a spend limit enforced in the loop stop it.

Attribution, or you cannot manage any of this

If all traffic goes through one untagged API key, you know your total spend and nothing else. You cannot tell which feature, tenant, or customer is expensive, so you cannot make any decision except panic.

Route calls through a gateway that tags every request with feature, tenant, and environment. Then the question stops being “why is the bill high” and becomes “the document-summary feature costs $0.40 per use and is used twice a month” — which is an answerable business question.

What to carry forward

Output costs several times input. Model the layer stack, measure cost per task, and remember that thinking tokens and agent steps multiply invisibly. Cache the stable prefix, right-size the model, cap the output, and put hard ceilings inside every loop. Tag everything.

And do the arithmetic before you build. Ten minutes with this formula will change more designs than any optimisation you apply afterwards.

That closes Part IV, and with it the fifteen mechanisms. One chapter remains, and it is the only one in the book that teaches no mechanism at all. You now have fifteen prices. What you do not have yet is the order in which to spend them.

RECALL

  1. Why does output cost several times more than input? Tie it to Chapter 4.
  2. Write the cost-per-task formula from memory.
  3. Size this before building: 2,000-token prompt, 4,000 tokens retrieved, 600 out, 3 steps, 5% retries, 20,000 tasks a month, $3/M in and $15/M out.
  4. Why are thinking tokens and agent loops the two most dangerous multipliers?
  5. Your bill tripled with no traffic change. Name the four likeliest causes.
  6. Why is per-feature attribution a prerequisite for cost control rather than a reporting nicety?