Tokens: The Unit of Everything
ONELINE
The model never sees your text. It sees a list of integers — and it bills you, limits you, and fails you in those integers, not in words.
Imagine hiring a translator who owns a box of about 100,000 pre-printed cards. Some cards hold a whole common word — the, running, import. Some hold a fragment — un, ing, ##bed. To say anything at all, they must build the sentence from cards they already own. They cannot write a new one.
That is a tokenizer. Everything strange about how models count, charge, and forget follows from it.
This is the first chapter because it is the unit the rest of the book is denominated in. Context windows are measured in tokens. Bills are measured in tokens. Latency is measured in tokens per second. If you are fuzzy here, every number later in the book stays fuzzy.
Where the cards come from
Someone had to print those phrase cards, and they did it once, before the translator was hired. That is the half of tokenization most explanations skip, and skipping it is what makes the rest feel arbitrary.
A tokenizer is trained by its own job, on its own corpus, before the model it serves exists. The job answers one question: which pieces of text are common enough to deserve a card of their own? The algorithm that answers it is almost always Byte Pair Encoding, and its procedure is surprisingly simple.
Start with every byte as its own token. Then repeat: find the most frequent adjacent pair in the corpus, and glue it into a single new token. Stop when you hit your target vocabulary size.
Watch it run on a tiny corpus, low lower lowest:
start l o w _ l o w e r _ l o w e s t
merge 1 ('l','o') -> lo w _ lo w e r _ lo w e s t
merge 2 ('lo','w') -> low _ low e r _ low e s t
merge 3 ('low','e') -> low _ lowe r _ lowe s t
Three merges in, low is one token. Words the corpus saw often become single cards. Words it rarely saw stay in pieces. That single fact — frequency decides what is atomic — explains most of what follows, and it was settled once, by a corpus you did not pick.
The job emits two artefacts: a vocabulary, every token paired with an integer id, and the ordered merge rules that built it. Both are then frozen. Only then is the model trained, on the ids that tokenizer produces — which is why a model and its tokenizer ship together and cannot be swapped independently.
ASIDE
This is also why you cannot teach a model a new word by sending it documents. A new token needs a new tokenizer, and a new tokenizer needs a new model.
Nothing on that diagram happens while you are talking to the model. Which brings us to the other half.
What actually happens to your text
At runtime the tokenizer learns nothing. It is a lookup: your text arrives, the frozen merge rules are applied in the order they were discovered, and out come ids from a vocabulary settled long before you typed. Send it a million of your own documents and the vocabulary does not move.
Take the cat sat. Text enters as characters and leaves as integers, and the path is short:
Those three ids are the model’s actual input. Everything after this chapter — the context window, the latency, the invoice — is denominated in them. The fourth step, turning ids into coordinates, is chapter 2’s subject.
Hold on to that shape. Build time against runtime is the same split you meet in chapter 7, where documents are embedded once offline and only the query is encoded per request, and again in chapter 8, where whatever you destroy at index time is gone at query time.
Why counting the r’s in “strawberry” can be difficult
This is the most-asked interview question about tokenization, and the answer is not “models are bad at counting.” It is that tokenization changes the granularity at which the model sees text.
Suppose the tokenizer represents strawberry as three tokens:
strawberry -> [st] [raw] [berry]
The model therefore receives three token ids, not ten character ids — something like:
[417, 892, 1534]
The characters have not disappeared. The tokenizer can still decode those three tokens back into strawberry, letter for letter. But the model’s immediate input representation is now the three larger pieces, and the three rs are distributed across them:
st -> 0 r's
raw -> 1 r
berry -> 2 r's
To answer “how many r’s?”, the model has to reason inside its tokens rather than count its input elements. It often gets there — spelling the word out piece by piece first is exactly that reconstruction — but it is extra work, and work that can fail.
The general form is worth memorising, because it predicts failures you have not met yet:
When a task depends on structure below the token level, the model must reconstruct that structure before it can solve the task.
Character counting. Exact string manipulation. Reversing strings. Spelling, rhyme, and syllable work. Arithmetic on long numbers, where digits group into tokens in ways that ignore place value. These are all the same difficulty wearing different clothes.
The three consequences you will actually feel
1. Cost
You are billed per token, input and output, usually at different rates. This makes token count a direct line item, not an implementation detail.
NAPKIN — What a support bot costs
Say each conversation sends a 2,000-token system prompt and 500 tokens of history, and gets back 300 tokens.
Input: 2,500 tokens. Output: 300 tokens.
At $3 per million input and $15 per million output:
(2500 / 1e6) x 3 = \$0.0075 plus (300 / 1e6) x 15 = \$0.0045
About $0.012 per conversation. At 100,000 conversations a month, that is $1,200 — and roughly two-thirds of it is the system prompt you send every single time.
Chapter 5 shows how to stop paying for that prompt repeatedly.
Notice what the arithmetic surfaced: the fixed part of the prompt dominated. That is typical, and invisible until you count tokens.
2. Limits
The context window is a token budget, not a character budget or a page budget. “Can I fit this contract in?” is a question only the tokenizer can answer.
Rules of thumb for English, useful for planning and useless for billing:
| From | To tokens | Notes |
|---|---|---|
| 1 word | ~1.3 tokens | Prose. Code runs higher. |
| 4 characters | ~1 token | The classic estimate. |
| 1 page | ~500–800 tokens | Depends heavily on formatting. |
WATCHOUT
These estimates are for capacity planning only. Never use len(text) / 4 in code that enforces a limit or computes a bill. Use the real tokenizer. The estimate is wrong in exactly the cases that hurt — code, JSON, non-English text, and long numbers — and it is wrong in the direction of undercounting, so you discover it as a production truncation or an invoice.
3. Capability, unevenly distributed
Tokenizers are trained on a corpus, and that corpus is not language-neutral. A tokenizer that saw mostly English spends few tokens on English and many on everything else.
The effect is a tax. The same meaning, expressed in a different language, costs more:
| Tokenizer generation | Chinese | Japanese | Hindi |
|---|---|---|---|
| Early (GPT-2 era, 50k vocab) | 2.5x | 3.0x | 6.0x |
| Mid (100k vocab) | 1.4x | 1.6x | 3.2x |
| Current (128k–200k vocab) | ~1.1–1.2x | ~1.2–1.3x | ~1.4–1.5x |
Read the trend, not the cells. Vocabularies grew from 32k to 128k and beyond largely to close this gap, and it has mostly closed. But “mostly” is doing work: a Hindi-language product on an older tokenizer paid up to six times more per sentence than its English equivalent, for identical meaning, and its effective context window was six times smaller.
ASIDE
Larger vocabulary is not free. It grows the embedding table and the output layer, which costs memory on every forward pass.
If you serve non-English users, tokenizer choice is a product decision, not an infrastructure one.
Special tokens, and the bug you will hit once
Alongside learned text tokens, models reserve a handful of special tokens for structure: start of sequence, end of sequence, padding, and — most importantly for you — chat role markers.
A chat conversation is not sent as a nice object. It is flattened into one string with delimiters:
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant
The model learned these markers during training. Get them wrong — wrong delimiter, missing role, hand-rolled formatting — and quality degrades in a way that looks like the model got dumber, because from the model’s side it did: it is now reading a shape it never saw in training.
WATCHOUT
Use the tokenizer’s own chat template rather than building the string yourself. Every format is model-family specific, and a format mismatch produces quiet quality loss rather than a loud error — the worst possible failure mode.
Tokens are not just text any more
Images, audio, and video enter the same way: as tokens.
An image is cut into patches — typically 14x14 pixels — and each patch becomes a visual token. A single image commonly costs a few hundred to a thousand tokens, depending on resolution. Audio is compressed into discrete units by a codec and fed in as a sequence. Video is handled as frames, so one second of video at one frame per second can cost about as much as one image.
The practical consequence: multimodal context is expensive in the same currency. A chat with twenty screenshots in it is not a light conversation. It is a long one, and it will hit your context ceiling long before the text does.
What to carry forward
Tokens are the unit of account for this entire field. When you cannot explain why something costs what it costs, or why it did not fit, or why the model cannot do a task that looks trivial, count tokens first.
The next chapter takes the integers this one produced and asks what happens when they become coordinates — because that is where meaning enters the picture.
RECALL
- Walk through BPE on
low lower lowest. After three merges, which sequences are single tokens, and why those — and when did those merges get decided? - Explain to a sceptical PM why the model miscounts letters in a word, without using the phrase “it’s bad at counting.”
- A user reports the model silently dropped the end of their document. Name two token-level causes.
- Your product launches in Hindi on a 50k-vocab tokenizer. What happens to your per-user cost and your effective context window, and by roughly how much?
- Why is
len(text) / 4acceptable for capacity planning but not for billing? - Which of these are token-boundary problems? Counting vowels; summarising a PDF; reversing a string; rhyming; retrieving a phone number.