The Context Budget

Lesson 06 / 16 · updated 2026-10-01 · 8 min


ONELINE

A context window is a budget, not a container. Bigger windows raised the ceiling and changed nothing about the discipline of deciding what deserves the space.

METAPHOR

A carry-on bag. The airline raised the size limit; you still cannot bring the house.

Part I was about a machine you do not control. You cannot change how attention scales or how much memory a KV cache eats.

What you do control, completely, is what goes into the context. That is the subject of Part II, and this chapter is the framing for all of it.

The temptation, and why it fails

Context windows grew from 4,000 tokens to a million and more. The obvious conclusion — stop being clever, put everything in — is wrong, for three separate reasons.

It is quadratically expensive. Chapter 3’s table: 32x the context is roughly 1,000x the attention compute, and past about 100k tokens attention is the majority of what you are paying for. You feel it as latency and as an invoice.

It caps your concurrency. Chapter 5: context is per-user memory. A system where every request carries 200k tokens serves a fraction of the users of one that carries 8k.

And accuracy degrades. This is the one that surprises people.

Context rot

Models get worse as the context grows, even well inside the stated limit.

Part of this is attention: with more tokens, the weight any single token receives is spread thinner. Part of it is training — models saw far more short sequences than million-token ones, so the long end is less well-practised.

The related, older finding is lost in the middle: information placed in the middle of a long prompt is recalled less reliably than the same information at the start or end. Frontier models handle this much better than early ones did, but the gradient has not vanished. It has flattened.

The practical rule survives regardless of model generation:

WATCHOUT

Put critical instructions and your best examples at the beginning and the end of the prompt. Put bulk reference material in the middle.

When you retrieve chunks, order them so the most relevant land first and last, not buried at position 14 of 20. A reranker that improves ordering (Chapter 9) helps twice: better chunks, and better placement of those chunks.

The deeper reframe is this: a large window is a resource with diminishing returns, not free space. Your job is to find the smallest high-signal set of tokens that still lets the model act correctly.

Budgeting, concretely

Treat the window like any other fixed resource: allocate it in advance.

Context budgetFlowchart with 6 labelled stages: Context window (the ceiling; System prompt 500 - 1,000; History 2,000 - 5,000; Retrieved elastic; Output reserve 1,000 - 4,000; Forget this one and the model truncates mid-answer. Connections: Context window (the ceiling leads to System prompt 500 - 1,000; Context window (the ceiling leads to History 2,000 - 5,000; Context window (the ceiling leads to Retrieved elastic; Context window (the ceiling leads to Output reserve 1,000 - 4,000; Output reserve 1,000 - 4,000 leads to Forget this one and the model truncates mid-answer.Context window (the ceiling)System prompt500 - 1,000History2,000 - 5,000RetrievedelasticOutput reserve1,000 - 4,000Forget this one and themodel truncates mid-answer
Every component gets an allocation. The output reserve is the one teams forget.
Component Typical budget Why
System prompt 500–1,000 tokens Core logic and persona. Longer is rarely better.
Conversation history 2,000–5,000 tokens The running state
Retrieved data 10k and up The elastic part, sized to the task
Output reserve 1,000–4,000 tokens Space for the answer and any reasoning

WATCHOUT

The output reserve is the allocation teams forget. The window is shared between input and output — fill 127k of a 128k window with input and the model has 1k left to answer in. On a reasoning model that also spends tokens thinking, you can starve it into truncating mid-sentence.

The symptom is answers that stop abruptly, and it gets misdiagnosed as a model quality problem.

Five techniques for staying inside the budget

When a single request cannot be trimmed enough — and especially once agents enter in Part III, where context accumulates turn after turn — these are the five moves.

Compaction. When the history gets long, summarise it and restart the loop with the summary plus the few most recent artefacts. Keep the load-bearing details — decisions made, constraints discovered, bugs still open — and drop redundant tool output. Tune for recall first (keep everything that matters), then trim for precision.

Just-in-time loading. Keep lightweight references in context — file paths, URLs, row ids — and fetch the full content only when a step needs it. This is how a person works in a codebase: you hold the file tree in your head and open the one file you need, rather than reading the whole repository first. The cost is latency on each fetch, so hybrids are common: preload the obvious, fetch the rest.

Structured note-taking. The agent writes progress notes to a file or store outside the window, then reads them back when relevant. This is what lets a task run far longer than the window it runs in. Chapter 12 is about the storage side of this.

Sub-agent isolation. Hand a focused sub-task to a sub-agent with its own clean window; it returns a 1–2k token summary. The intermediate detail — twenty search results, a long file read — never touches the coordinator’s context. This is a context-management argument for multi-agent systems, entirely separate from any parallelism benefit.

System prompt calibration. Aim for the zone where the prompt is specific enough to be reliable but general enough not to be brittle. Use clear sections. Resist the urge to patch every failure by appending another sentence — that is how 500-token prompts become 6,000-token prompts that are worse.

ASIDE

Prompt engineering writes one good instruction. Context engineering curates every token the model sees, every turn. The second is the production discipline.

Where the money actually is

Chapter 5 introduced prompt caching. It belongs here too, because it changes what “expensive” means.

The costly part of a long prompt is usually the stable part — a system prompt, tool definitions, few-shot examples — sent identically on every request. Cached, that content costs a fraction and skips most of prefill.

Which yields a layout rule that is worth more than most prompt tweaks:

[ stable ]  system prompt, tool defs      <- cacheable
[ stable ]  long reference docs that rarely change
[ varies ]  retrieved chunks for this query
[ varies ]  conversation history
[ varies ]  the user's message            <- never cached

Stable first, variable last. Anything variable near the top breaks the prefix match and forfeits the cache for everything after it.

Does retrieval still matter with a million tokens?

Yes, and it is worth being able to defend this crisply, because it is a standard interview question.

  • Cost and latency. Re-reading a million tokens on every query is enormously more expensive than retrieving the five relevant chunks, even with caching.
  • Freshness. Retrieval can hit a live API or a database updated a second ago. A context window holds a snapshot you pasted in.
  • Scale. Enterprise corpora are terabytes. They do not fit in any window, and never will.

The honest nuance: for a genuinely small, stable corpus — under about 50k tokens — skipping the vector database and putting everything in a cached prompt is often the better engineering decision. Simpler, no index to maintain, no chunking decisions, no embedding model migrations.

Know which regime you are in. Building a retrieval pipeline for a 30-page handbook is over-engineering; putting a 4-terabyte document store in a prompt is not an option.

What to carry forward

Context is a budget with diminishing returns. Allocate it, reserve room for the output, put critical material at the edges, order it so the stable parts cache, and reach for the five techniques when it overflows.

The next three chapters are about the hardest line item in that budget: how you choose which few thousand tokens of a huge corpus deserve the space.

RECALL

  1. Give three reasons not to fill a million-token window even when you can.
  2. What is context rot, and how does it differ from the lost-in-the-middle effect?
  3. A model starts truncating answers mid-sentence on long documents. What is the likely budget error?
  4. Reorder this prompt for prompt caching: user question, retrieved chunks, system prompt, tool definitions. Explain the ordering.
  5. Name the five context-management techniques and the situation each fits.
  6. A colleague says long context made RAG obsolete. Give the three-part answer, then say when they are actually right.