The Generation Loop: Prefill and Decode
ONELINE
Generation is two different machines wearing one coat: a parallel, compute-bound read, followed by a serial, memory-bound write. Almost every latency complaint is really a question about which half you are in.
You send a prompt. Tokens come back. It looks like one operation, and it is billed like one operation, so it is easy to model it as one operation.
It is not. It is two, with opposite performance characteristics, and conflating them is the most common reason teams optimise the wrong thing for a quarter.
Phase one: prefill
The model reads your entire prompt in parallel. Every token is processed at once, attention is computed across all of them, and the keys and values for every position are written into the cache.
This is a big, dense matrix operation — exactly what a GPU is built for. It is compute-bound: the arithmetic units are saturated and the limit is how many FLOPs the chip can do.
Prefill happens once, and its cost scales with how long your prompt is — more than linearly, once the context is long enough for Chapter 3’s attention term to dominate. This is what you are paying for when you pay for input tokens.
Phase two: decode
Now the model writes. And it can only write one token at a time, because each token depends on the one before it. Token 50 cannot be computed until token 49 exists.
Each step is one full forward pass through the model to produce exactly one token.
Here is the counter-intuitive part. That pass has to load the model’s weights out of GPU memory — tens of gigabytes — to compute a single token. The actual arithmetic for one token is trivial. The GPU spends nearly all its time waiting for memory.
Decode is memory-bandwidth-bound. The compute units sit largely idle.
ASIDE
This is why generation speed tracks memory bandwidth far more closely than it tracks raw FLOPs when comparing accelerators.
Why this distinction earns its own chapter
Because the two phases respond to completely different fixes.
| Prefill | Decode | |
|---|---|---|
| Parallel? | Yes, all tokens at once | No, strictly one at a time |
| Bottleneck | Compute (GPU cores) | Memory bandwidth |
| Scales with | Prompt length | Output length |
| Happens | Once | Once per output token |
| Metric | Time to first token | Tokens per second |
| Faster by | Shorter prompts, caching | Bigger batches, smaller weights |
Two metrics, measured separately, because they answer different questions:
Time to first token (TTFT) — how long until something appears. This is network time, plus queue time, plus prefill. Dominated by prompt length. For interactive chat you want under 500 ms; for voice, under 200 ms.
Tokens per second (TPS) — how fast text streams once it starts. Dominated by memory bandwidth and batch size.
WATCHOUT
“The model is slow” is not a diagnosis. Ask which number is bad.
Slow TTFT with fine TPS means your prompt is too long or your queue is deep — look at prefill, caching, and prompt size.
Fine TTFT with slow TPS means you are memory-bound — look at batching, quantisation, and model size.
Optimising the wrong one is wasted work. A team that shrinks its model to fix a TTFT problem caused by a 20k-token system prompt will make quality worse and latency barely better.
The batching insight that pays for itself
Decode is memory-bound: most of the time goes to loading weights, and the arithmetic is nearly free.
So what happens if you process sixteen users’ next tokens in the same pass?
You load the weights once — the expensive part — and do sixteen times the arithmetic, which was never the constraint. Throughput goes up almost 16x. Each individual user’s tokens-per-second barely moves.
NAPKIN — Why serving is all about batch size
Suppose one decode step takes 20 ms, of which 19 ms is loading weights and 1 ms is arithmetic.
Batch of 1: 20 ms, produces 1 token → 50 tokens/sec total.
Batch of 16: 19 ms loading (unchanged) + 16 ms arithmetic = 35 ms, produces 16 tokens → about 457 tokens/sec total.
Nine times the throughput on the same hardware. Per-user speed dropped from 50 to about 29 tokens/sec — still faster than anyone reads.
This is why every serving stack fights for larger batches, and why the thing that limits batch size limits your economics. That thing is the KV cache, which is the next chapter.
Sampling: how the next token is chosen
The model does not output a token. It outputs a probability over every token in the vocabulary. Something must choose.
- Greedy — always take the most likely token. Repetitive, and not actually deterministic in practice across batch sizes and hardware.
- Temperature — flatten or sharpen the distribution before sampling. Low temperature is conservative, high is creative. Zero approximates greedy.
- Top-k / top-p — restrict the choice to the k most likely tokens, or to the smallest set whose probabilities sum to p. Cuts the long tail of nonsense without capping creativity.
The practical guidance is short: for extraction, classification, and structured output, use a low temperature. For open-ended writing, raise it. And do not expect temperature zero to give you reproducibility — it reduces variance, it does not eliminate it.
Streaming is a latency illusion, and a good one
Because decode produces tokens one at a time anyway, you can send each one as it appears rather than waiting for the whole response.
Total time is unchanged. But the user sees output after the first token instead of the last, which is the difference between 400 ms and 8 seconds of apparent latency.
Almost every user-facing LLM feature should stream. It is the cheapest latency win available, and it costs you nothing but a slightly more involved client.
WATCHOUT
Streaming complicates anything that needs the complete output before acting — JSON validation, guardrails, tool-call parsing. You either buffer (losing the benefit) or validate incrementally (more code). Decide this before you build, not after. Chapter 14 comes back to it.
What to carry forward
Two phases, two bottlenecks, two metrics. Prefill is compute-bound and scales with your prompt; decode is memory-bound and scales with your output. Batching fixes decode throughput. And the thing standing between you and bigger batches is the subject of the next chapter.
RECALL
- Why is prefill compute-bound while decode is memory-bound?
- A user reports “it takes forever to start, then it’s fine.” Which metric is bad, and what are your first two moves?
- Explain why batching helps decode throughput enormously but barely helps per-user speed.
- Why can decode not be parallelised across tokens the way prefill can?
- When would you not stream a response?
- Your team sets temperature to 0 and expects identical outputs across runs. What will they observe, and why?