Embeddings: Meaning as Geometry
ONELINE
An embedding turns a model’s learned representation of meaning into coordinates, which makes similarity something you can compute — and quietly imports every assumption of the space you projected into.
Chapter 1 left you with integers. Integers are not meaning — 1820 is no closer to 1821 than to 9000. Something has to bridge that gap.
That something is an embedding: a list of numbers, typically 256 to 4,096 of them, positioned so that things that mean similar things sit near each other.
Think of an ordinary library. Books are shelved by subject, so proximity carries information: if you find one book on Roman history, its neighbours are probably also about Roman history. A shelf is one-dimensional, and a floor plan is two. An embedding space is the same idea with a thousand dimensions, which is enough room for “about Rome”, “written for children”, “in French”, and “argumentative in tone” to each be their own direction.
No individual dimension has a meaning you can name. Dimension 412 is not “formality”, and dimension 87 is not “about Rome”. Meaning emerges from the combination of dimensions, not from any single one of them.
Why this is the enabling trick
Once meaning is geometry, similarity is something you can compute. Two representations can be compared with arithmetic, rather than requiring a model to reason about their relationship from scratch.
That one property supports a surprising amount: retrieval, clustering, deduplication, recommendation, classification. One representation, many uses. Part II is built on it.
That list of numbers has a name: a vector. A vector is the numerical representation of a point in the embedding space, and from here on the two words mean the same thing.
The classic demonstration is that some relationships show up as directions you can do algebra on:
vector("king") - vector("man") + vector("woman")
~= vector("queen")
This is a real property of trained spaces and it is worth seeing once, for intuition. Do not mistake it for a general rule: the neat arithmetic works on simple analogies and breaks down on anything more complex.
Static versus contextual: the distinction that matters
The first generation of embeddings gave each word one fixed vector. Word2Vec, GloVe, and FastText all produce these static word representations. One word, one point, forever.
That breaks as soon as a word means different things in different sentences. Take “I sat on the river bank” against “I opened a bank account”. In the first, bank is a riverbank; in the second, a financial institution.
A static embedding assigns bank one fixed vector regardless of the sentence around it, so it has to serve both meanings with the same point — a compromise that represents neither accurately.
Modern language-model embeddings are contextual: the representation of a token depends on the context in which it appears.
# Static: one vector, always
bank -> [0.1, 0.3, ...]
# Contextual: the surrounding tokens change it
"river bank" -> bank -> [0.1, 0.3, ...] # geography
"bank account" -> bank -> [0.5, 0.2, ...] # finance
The mechanism that makes this possible is attention, which is Chapter 3.
From token id to contextual representation
Inside a language model, the thing doing the first step is the embedding layer, and it is simpler than its reputation: a table with one row per vocabulary entry. Look up the row, get the vector.
Suppose token id 6201 represents bank:
The lookup is pure table access, and that matters more than it sounds.
The word bank is the same token id in both sentences — so it starts as the same vector in both. Nothing in the embedding layer can tell a riverbank from a savings account. That distinction is added later, as the Transformer layers incorporate the surrounding context.
That is the progression to carry into Chapter 3:
token id
|
embedding vector
|
Transformer layers
|
contextual representation
Chapter 3 looks at the mechanism at the heart of that transformation: attention.
NAPKIN — What the embedding table costs
One row per vocabulary entry, one column per model dimension.
A 100,000-token vocabulary at 4,096 dimensions in 2-byte floats:
100,000 x 4,096 x 2 = 819,200,000 bytes
About 820 MB, before a single layer of the model exists. Double it if the output layer has its own copy rather than reusing this one.
This is the hidden half of Chapter 1’s point that a larger vocabulary is not free. Doubling the vocabulary to 200k adds roughly another 820 MB of weights that must be loaded and kept resident.
ASIDE
The embedding table is usually the single largest weight matrix in a small model, and nearly irrelevant in a large one. The rest of the network grows faster than the vocabulary does.
A note on the word “embedding”
The term gets used for two related things, and conflating them causes real confusion later in this book.
| This chapter | Part II | |
|---|---|---|
| What is embedded | One token | A chunk of text |
| Produced by | The embedding layer inside the model | A dedicated embedding model |
| Purpose | Give the Transformer something to compute on | Make a corpus searchable |
| Depends on context | Yes, after attention | No — encoded before the query exists |
Same word, same underlying idea — meaning as coordinates — but a different job. Here it is the layer that turns discrete token ids into something a neural network can work with. Text embeddings for search return in Chapter 7.
What to carry forward
Tokens gave the model discrete symbols. Embeddings turn those symbols into vectors a neural network can compute with. The embedding layer supplies the starting representation, identical every time the token appears; everything that makes a representation situation-specific is added by the layers above it.
Next, those layers — and the reason long context costs what it does.
RECALL
- Why is a token id not already a representation of meaning? What does the embedding layer do about it?
- Why does a static word embedding fail on “river bank” versus “bank account”, and what changed to fix it?
king - man + woman ~= queenis real. Why is it a bad thing to build on?- Distinguish the two things called an “embedding” in this book. Which one does a Transformer need in order to run at all?