The KV Cache: The Memory That Runs the Bill
ONELINE
The cache that makes generation affordable is also the hard ceiling on how many users you can serve at once. Learn to see it and serving costs stop being mysterious.
Chapter 4 ended on a cliff: batching is how you get throughput, and something limits your batch size. This chapter is about that something.
The problem it solves
Decode generates one token per forward pass. To generate token 101, attention needs the keys and values for tokens 1 through 100.
The naive approach recomputes them every step. That means generating a 500-token answer does the work of processing the prompt 500 times over. It is quadratic, and it is unusable.
The fix is obvious once stated: compute each token’s key and value once, then keep them. That store is the KV cache. Every serving system uses it. Without it, generation as a product would not exist.
So far, so good. Now the bill.
The cache is per-user, and it grows
Model weights are loaded once and shared by everyone. The KV cache is not. Every concurrent user has their own, and it grows with every token in their conversation.
The size is fully determined by the architecture and the context length:
2 (K and V)
x layers
x context length
x KV heads
x head dimension
x bytes per value
NAPKIN — One user at 128k context
For a 70B-class model with 80 layers, 8 KV heads, head dimension 128, at 16-bit precision, holding a full 128,000-token context:
2 x 80 x 128,000 x 8 x 128 x 2 bytes ≈ 42 GB
Forty-two gigabytes. For one user’s conversation — on top of roughly 140 GB of model weights.
On a 8x80GB node (640 GB total), weights leave about 500 GB. That is room for roughly twelve such users. Twelve.
That number is the whole chapter. Your concurrency is not set by how clever your autoscaler is. It is set by arithmetic you can do before you build.
Where the 8x already went
Notice the calculation used 8 KV heads, not the 64 attention heads such a model would typically have. That is not a typo — it is the single biggest optimisation in modern serving, and it is already priced in.
In original multi-head attention, every attention head kept its own keys and values. With 64 heads, that same calculation gives about 344 GB per user, which is more than a whole node. Long context would simply be impossible.
Grouped-query attention (GQA) lets several query heads share one set of keys and values. Sixty-four query heads sharing eight KV groups gives an 8x reduction in cache size — 344 GB becomes 42 GB — at a quality cost small enough that it is now near-universal.
ASIDE
Multi-query attention is the extreme version: all query heads share a single KV head. Maximum saving, more quality loss. GQA is the compromise that won.
This is worth internalising because it reframes what GQA is for. It is not a quality technique. It is the thing that made long context economically possible, and you can see exactly how much it bought by dividing 344 by 42.
Fragmentation, and the fix that doubled the industry’s throughput
Early serving systems allocated each request a contiguous block of memory sized for the maximum possible output length, because they could not know in advance how long the answer would be.
A request that might generate 2,000 tokens but actually generates 100 wasted 95% of its allocation. Measured utilisation in real systems was often around 20–40%.
PagedAttention borrowed the fix from operating system virtual memory: break the cache into fixed-size blocks, allocate them on demand, and keep a lookup table mapping a sequence’s logical positions to wherever its blocks physically live. Blocks need not be contiguous.
Waste drops to near zero, and batch sizes roughly double or better on the same hardware. It also makes sharing cheap — several requests with the same prompt prefix can point at the same physical blocks instead of holding duplicates.
You will not implement this. You should know it is why modern serving stacks are several times more efficient than a naive implementation, and it is a strong reason to use one rather than rolling your own.
Prompt caching: the same idea, exposed to you
Here is where this becomes a decision you actually make.
If the KV cache for a prefix can be kept and reused, then a long system prompt sent on every request does not have to be re-processed every time. Its keys and values are already computed. Prefill skips straight past it.
Providers expose this as prompt caching, and the economics are dramatic: cached input tokens typically cost a fraction of normal input tokens, and TTFT drops substantially because prefill has far less work to do.
NAPKIN — Revisiting the Chapter 1 support bot
Chapter 1 found a support bot costing about $1,200 a month, roughly two-thirds of it the same 2,000-token system prompt sent 100,000 times.
That prompt is identical every request. With prompt caching at roughly a tenth of the input price, that $600-odd of repeated prefix drops to around $60.
Total falls from about $1,200 to about $660 — a 45% cut from one architectural choice, with no quality change whatsoever.
WATCHOUT
Caching keys on an exact prefix match. Put anything variable near the front of your prompt — a timestamp, a user id, a session token — and every request gets a different prefix, so nothing ever hits the cache.
Order matters: stable content first (system instructions, tool definitions, few-shot examples), variable content last (user message, retrieved chunks). Teams routinely enable prompt caching, change nothing about prompt order, see no saving, and conclude the feature does not work.
The tension you now own
Every serving decision comes back to one pool of memory:
- Longer context means a bigger cache per user, so fewer concurrent users.
- More concurrent users means less context each.
- Bigger model means less memory left for cache, so both get worse.
- GQA and quantisation shrink the cache, buying back both.
There is no configuration that escapes this. There is only the trade you choose, and you can compute it in advance.
What to carry forward
The KV cache is where the abstract quadratic cost of Chapter 3 becomes a number on an invoice. Per user, grows with context, capped by GPU memory. GQA bought an 8x reduction, PagedAttention removed the waste, and prompt caching hands you the lever directly — if you order your prompt to earn it.
That closes Part I. You now have the mechanism behind every cost and limit in the rest of the book. Part II turns to the thing you actually control: what goes into the context in the first place.
RECALL
- Why does the KV cache exist? What is the cost of not having one?
- Compute the cache for one user: 60 layers, 8 KV heads, head dim 128, 32k context, 2 bytes per value. Then say how many such users fit in 400 GB.
- GQA cuts the cache by 8x on a 64-head model. What exactly is being shared, and what is given up?
- Explain PagedAttention by analogy to operating system memory. What waste does it remove?
- Your team enabled prompt caching and saw no cost change. Name the most likely cause and the fix.
- Why does serving a larger model reduce the number of concurrent users by more than the weight increase alone suggests?