Attention: Why Context Is Expensive

Lesson 03 / 16 · updated 2026-10-01 · 13 min


ONELINE

Every token looks at every token before it. That one sentence explains the context limit, the shape of your latency curve, and most of the optimisation industry.

METAPHOR

A meeting where, before anyone may speak, everyone must shake hands with everyone else.

Chapter 2 stopped one step short. A token id becomes a vector by table lookup, and then Transformer layers turn that identical starting vector into something that knows which sentence it is in. This chapter opens the layer.

Here is the problem it solves. In the sentence “The trophy did not fit in the suitcase because it was too big”, what does it refer to? You resolved that instantly, using the rest of the sentence. A model processing tokens in isolation cannot. It needs a mechanism for every position to look at every other position and pull in what is relevant.

Query, key, value

Attention borrows the vocabulary of search, and the analogy is exact enough to be worth taking literally.

Part The question it answers Search analogy
Query What am I looking for? Your search box text
Key What do I contain? A document’s index entry
Value What do I contribute? The document’s actual content

Each token produces all three, by multiplying its vector by three learned weight matrices. Then for each token: compare its Query against every Key to get a relevance score, softmax those scores into weights that sum to one, and take the weighted blend of all the Values.

Attention qkvFlowchart with 8 labelled stages: Token: 'it'; Query: what am I looking for?; Compare against every other token's Key; All tokens' Keys: what do I contain?; Attention weights (soft, sums to 1; Weighted blend of Values; All tokens' Values: what do I contribute?; New vector for 'it', now aware of context. Connections: Query: what am I looking for? leads to Compare against every other token's Key; Compare against every other token's Key leads to Attention weights (soft, sums to 1; Attention weights (soft, sums to 1 leads to Weighted blend of Values; Weighted blend of Values leads to New vector for 'it', now aware of context.Token: 'it'Query:what am I looking for?Compare against everyother token's KeyAll tokens' Keys:what do I contain?Attention weights(soft, sums to 1)Weighted blend of ValuesAll tokens' Values:what do I contribute?New vector for 'it',now aware of context
Attention for a single token. Every token runs this, against every token before it, in every layer.

For it, the Query roughly encodes “I am a pronoun, I need a noun”. The Key for trophy matches it well, the Key for because does not. So the output vector for it ends up mostly made of trophy’s Value. The pronoun has been resolved, geometrically.

ASIDE

“Scaled” dot-product attention divides scores by the square root of the dimension. Without it, large dimensions produce huge scores and softmax saturates.

Two properties are worth pausing on.

It is soft. The model does not pick one token to attend to. It spreads weight across all of them. That is what makes it differentiable, and therefore learnable. It also means the weights are a fixed budget: they sum to one, so a longer context divides the same attention among more tokens. Chapter 6 collects the bill for that.

It is direct. Older architectures passed information along a chain, one step at a time, so distant words communicated through a long game of telephone. In attention, position 1 and position 4,000 are one hop apart. This is why transformers handle long-range dependencies so much better — and, as you are about to see, why they cost what they do.

Only backwards

The one-sentence version of attention says every token before it. Two things are hiding in that phrase, and the first one is load-bearing.

In a model that generates text, a token may look at itself and at the tokens before it, and at nothing after. Position 50 sees positions 1 through 50. It may not see position 51, because during training position 51 was the answer, and a model allowed to read the answer learns nothing.

This is causal masking, and it is one line of arithmetic: scores for later positions are set to minus infinity before the softmax, so their weights come out as zero.

             the  trophy  did  fit
  the         x     .      .    .
  trophy      x     x      .    .
  did         x     x      x    .
  fit         x     x      x    x

  x = may look here    . = masked, weight zero

Two consequences, both of which you meet again.

The comparison count halves — it is n squared over two, not n squared. That changes the constant and leaves the shape alone, which is why the rest of this chapter still says quadratic.

More importantly: a token’s Key and Value never change once computed. Nothing that arrives later is permitted to affect them. That single fact is what makes the KV cache of Chapter 5 possible at all, because a cache is only safe when the thing you cached cannot go stale.

Where word order comes from

The second thing hiding in it. Read the mechanism again and notice what is absent: position. Queries and keys are compared pairwise, and a softmax does not care what order its inputs arrive in. Shuffle the words of a sentence and plain attention computes the same set of blends.

So order has to be injected. Early models added a position vector to each token embedding before the first layer. Current ones rotate the Query and Key vectors by an angle that depends on position, which makes a score depend on how far apart two tokens are rather than where they sit absolutely.

ASIDE

The rotation trick is RoPE, rotary position embedding. It generalises to longer inputs better than an added vector, which is why it won.

You will never tune this. It is worth knowing for one reason: “supports 128k context” is partly a claim about position encoding, not only about memory. A model stretched well past the length it was trained on is being asked about positions it has little practice representing — one of the reasons quality sags inside the advertised window.

One layer, and then eighty more

Everything so far happens inside a single Transformer layer, and a layer is smaller than its reputation. It has two steps:

  1. Attention — the step you just read. Tokens look at each other and mix. This is the only step in the entire model where positions talk.
  2. A feed-forward network — each position is then processed on its own, with no reference to its neighbours at all.
Transformer layerFlowchart with 4 labelled stages: Vectors in - one per token; Attention - the only step where positions talk to each other; Feed-forward - each position processed alone, neighbours ignored; Vectors out - same count, revised not replaced. Connections: Vectors in - one per token leads to Attention - the only step where positions talk to each other; Attention - the only step where positions talk to each other leads to Feed-forward - each position processed alone, neighbours ignored; Feed-forward - each position processed alone, neighbours ignored leads to Vectors out - same count, revised not replaced.and again, 80 timesVectors in - one per tokenAttention - the only stepwhere positions talk to each otherFeed-forward - each positionprocessed alone, neighbours ignoredVectors out - same count,revised not replaced
Inside one layer. Only the attention step lets positions see each other.

Each step adds its result back into the running representation rather than replacing it, so a token’s vector is revised layer after layer rather than rebuilt. A 70B-class model stacks about 80 such layers.

Hold on to that number. It multiplies everything in the next section: every layer runs its own attention, over the whole sequence, with its own weights.

The bill for that directness

Read the mechanism again: every token compares against every token before it.

With 1,000 tokens that is about half a million comparisons. With 10,000 tokens it is about 50 million. Ten times the input, a hundred times the work — and all of it repeated in all 80 layers.

This is the quadratic cost of attention, and it is the single most important cost fact in this book:

Context Attention compute KV cache (70B, no GQA)
4K baseline 10.7 GB
8K 4x 21.5 GB
32K 64x 86 GB
128K 1024x 344 GB

Sit with the last row. Going from a 4,000-token context to a 128,000-token context — 32 times more text — costs about 1,000 times more attention compute, and needs 344 GB just to hold the intermediate state. That is more memory than the largest single GPUs of the era, for one conversation.

That last column is the original multi-head arrangement, one set of keys and values per head. Chapter 5 shows what the industry did to it — the 344 GB becomes 42 GB — without changing the column’s shape, which stays linear in context and private to each user.

NAPKIN — How much slower is a long prompt, really

Attention is quadratic. The rest of the model is not: every other matrix in a layer processes each token independently of the others, so its cost is flat per token. Prefill pays both, in billions of operations per token:

  context    flat   attention   share
    4,000     140        5        4%
   32,000     140       43       24%
  128,000     140      172       55%

Eight times the text, from 4k to 32k, therefore costs about 10x the prefill. Not 8x, and not 64x. If the short request’s prefill took 200 ms, budget about 2 s for the long one, not 12 s.

The 64x is real, but it applies to the attention term alone, and that term stays small until the context is large. The two cross at roughly 100,000 tokens; past there attention is the majority of your bill, and every further doubling of context quadruples it.

Which is why “just put the whole document in the prompt” stops being a good answer somewhere between 50k and 100k tokens — and why Chapter 7 exists.

Notice the KV cache column scales differently: linearly, not quadratically. It gets its own chapter (Chapter 5) because it is the thing that actually caps how many users you can serve at once.

Multiple heads

One more wrinkle. Real models do not run attention once per layer; they run it several times in parallel with different learned weights, then concatenate the results. These are attention heads.

The intuition: one head can specialise in syntax (which verb goes with which subject), another in coreference (what it refers to), another in position. Forcing one attention pattern to serve every relationship at once would be a bottleneck. Twelve to a hundred heads per layer is typical.

You rarely tune this. It matters mostly because it explains why the KV cache is as large as it is — every head keeps its own keys and values.

How the industry fought back

The quadratic wall produced an enormous amount of engineering. You do not need to implement any of it, but you should recognise the three families, because they are what model cards and vendor claims are talking about.

Sparse attention — do not let every token see every token. Give each one a local window, plus a few global tokens that everyone can see. Cost drops toward linear. The catch is that you have decided in advance which long-range links matter, and sometimes you are wrong.

Linear attention — algebraically restructure the computation to avoid ever forming the full n-by-n score matrix. Genuinely linear. Historically paid for it in quality, though the gap has narrowed.

Flash Attention — the one that actually changed production systems. It computes exact standard attention, but reorganises the memory access so the big intermediate matrix is never written to slow GPU memory. Same maths, same results, several times faster and far less memory.

WATCHOUT

Flash Attention reduces memory movement, not arithmetic. Attention is still quadratic in compute with Flash — people routinely misstate this in interviews. What Flash removes is the memory blow-up, which is what used to make long context impossible outright rather than merely expensive.

The distinction matters when you predict behaviour. With Flash, doubling context still roughly quadruples attention compute. It just no longer runs you out of memory first.

What this buys you as a designer

Three durable consequences:

  1. Prompt length is not free, and not linear. Halving a system prompt from 8k to 4k tokens saves about three-quarters of its attention cost. This is why Chapter 6 treats context as a budget.

  2. Long context and high concurrency are in direct tension. Memory spent on one user’s context is memory not available for another user’s. You choose.

  3. Input and output cost differently. Processing a prompt is one parallel pass; generating is a serial loop. They have different bottlenecks and different fixes — which is exactly where the next chapter goes.

RECALL

  1. Explain query, key, and value using the search analogy, then say what the softmax step is for.
  2. A request grows from 8k to 32k tokens of context. By roughly what factor does the attention term grow, and by roughly what factor does the whole prefill grow? Why are those two answers different?
  3. Why are transformers better than chained architectures at long-range dependencies?
  4. A colleague says “we enabled Flash Attention, so context length is linear now.” Correct them precisely.
  5. A token’s Key and Value never change after they are first computed. Why not, and which later optimisation is built entirely on that fact?
  6. Shuffle the words of a sentence. Why does plain attention not notice, and what is added to the model so that it does?
  7. Why do models use many attention heads instead of one?