The Retrieval Gap
ONELINE
RAG fails in exactly three places, and naming which one you are in tells you which fix to reach for. Reaching for the wrong one is the most commonly wasted quarter in this field.
The idea behind retrieval-augmented generation is almost embarrassingly simple. The model does not know your private data. So look the relevant part up, put it in the prompt, and ask the question.
Simple to state, and it demos beautifully. It also fails in production constantly, and teams respond by swapping the model — which almost never helps, because the model is rarely the problem.
Why look things up instead of training it in?
The alternative is fine-tuning: bake the knowledge into the weights. The comparison is worth memorising:
| Fine-tuning | Retrieval | |
|---|---|---|
| Where knowledge lives | In the weights | In the context |
| Updating it | Retrain | Update a row |
| Attribution | None — a black box | Explicit citations |
| Removing a document | Very hard | Delete it |
| Access control | Nearly impossible per-user | Filter at query time |
The rule of thumb: fine-tune for form, retrieve for fact. Fine-tuning teaches style, tone, format, and task-shape. Retrieval supplies knowledge.
The rows about attribution and deletion are not soft benefits. In a regulated setting, “show me the source” and “delete this customer’s data” are legal requirements, and only one column can satisfy them.
How vector search actually works
“Look the relevant part up” is carrying a lot of weight in that sentence. Here is the architecture almost every system uses. It is called a bi-encoder, because the query and the documents pass through the encoder separately.
Documents are encoded ahead of time and stored. At query time you encode only the query, then find the nearest stored vectors. Chapter 2’s geometry is what makes “nearest” mean anything.
This shape dominates because it moves nearly all the work off the hot path. A million documents can be encoded overnight; a query needs one encoder pass and a nearest-neighbour lookup.
But look closely at what was given up. The document was encoded without ever seeing the query. The encoder had to guess, in advance, every question that document might answer, and compress that into one point.
Hold on to that sentence. It is the root of the first two gaps below, and the reason Chapter 9 puts a second, slower stage behind this one.
One choice is worth making deliberately. A purpose-built embedding model beats a language model you pooled yourself, essentially always. Embedding models are trained with a retrieval objective — pull matching query and document pairs together, push mismatched ones apart. A general-purpose model was never asked to make that arrangement true.
Measuring closeness
“Nearest” needs a definition. Three metrics show up, and one of them matters.
| Metric | Measures | Use when |
|---|---|---|
| Cosine similarity | Angle between vectors | Default for text. Ignores length. |
| Dot product | Angle and magnitude | Vectors already normalised, or length is meaningful |
| Euclidean distance | Straight-line gap | Rare for text; common for images |
Cosine is the default because document length should not make a document more or less relevant, and raw magnitude tends to track length. Cosine throws magnitude away and keeps direction, which is where the meaning is.
NAPKIN — Why normalised vectors make dot product free
For vectors already scaled to length 1, cosine similarity is the dot product — the denominator is 1 x 1.
So most vector databases normalise on write, then use dot product on read. Same ranking, one fewer square root per comparison, across billions of comparisons.
If you ever see a system storing normalised vectors but computing full cosine, you have found free performance.
The three gaps
When a RAG system gives a bad answer, exactly one of three things went wrong. Diagnosing which is the entire skill.
Gap 1: Semantic mismatch
The answer is in your corpus, but the query and the document do not look alike to the retriever.
The user asks about “fast cars”. The document says “the Porsche 911 accelerates from 0-60 in 3.5 seconds” and never uses the word fast. Or — more commonly in technical corpora — the user searches NVIDIA_VISIBLE_DEVICES and dense retrieval, which works on meaning rather than exact strings, shrugs at an identifier it has no semantic intuition for.
Fixes: hybrid search, which adds keyword matching alongside semantics (Chapter 9); query rewriting, where the model rephrases or expands the question before searching.
Gap 2: Ranking failure
The document was retrieved. It came back at position 47. You passed the top 10 to the model.
This is a precision problem, not a recall problem, and it is the one most often misdiagnosed — teams see “the right answer wasn’t in the prompt” and conclude the search is broken, when the search actually found it.
Fixes: reranking, which re-scores candidates with a slower, more accurate model (Chapter 9); retrieving more candidates before filtering.
Gap 3: Lost in the middle
The right chunk was in the prompt, in the top five. The model still ignored it.
This is Chapter 6’s problem: attention spread thin across a long context, relevant material buried in the middle, drowned out by nine other chunks of plausible-looking noise.
Fixes: send fewer, better chunks; reorder so the strongest land first and last; compress chunks before inserting them.
WATCHOUT
You cannot fix a gap you have not identified, and all three present identically to the user as “the answer was wrong.”
Instrument the pipeline so you can tell them apart. For a sample of failures, check: was the right document in the corpus at all? Did it appear anywhere in the retrieved candidates? What rank? Was it in the final prompt?
Those four checks split every failure into gap 1, gap 2, gap 3, or “we never had the document” — and each has a different owner and a different fix. Chapter 13 turns this into a standing measurement.
The four shapes of RAG
Systems are usually described by how much agency the retrieval step has. The progression is worth knowing because it is a ladder, and most teams should climb it slowly.
Naive RAG. Query → vector search → top-k → generate. This is the demo. It is also the version that produces all three gaps at once, and it should not be the target architecture for anything in production.
Advanced RAG. Query rewriting → hybrid search → rerank → generate. Multi-stage, still a straight line. This closes gaps 1 and 2 and is the sensible default for most production systems. Chapters 8 and 9 are mostly about building this well.
Agentic RAG. The model decides what to search for, evaluates what came back, and searches again if it is not satisfied. A loop rather than a pipeline. Far more capable on hard questions, and considerably more expensive and harder to predict — every extra loop is another full request. Part III is about loops.
GraphRAG. Extract entities and relationships into a knowledge graph, then traverse it. The distinctive win is aggregative questions — “what themes recur across these 200 incident reports?” — which chunk-based retrieval fundamentally cannot answer, because no single chunk contains the answer. The cost is a substantial extraction pipeline.
ASIDE
Do not start at agentic or graph. Build the advanced pipeline, measure your three gaps, and let the measurements tell you whether you need more machinery.
Choosing between them
| Your situation | Start here |
|---|---|
| Corpus under ~50k tokens, fairly stable | No retrieval — cached long context (Chapter 6) |
| Standard question answering over documents | Advanced RAG: hybrid + rerank |
| Questions needing several lookups to answer | Agentic RAG |
| “Summarise across all documents” questions | GraphRAG |
| Corpus over 100k tokens and changing hourly | Advanced RAG, no question |
What to carry forward
RAG is not one thing that works or does not. It is a pipeline with three identified failure points, each with its own fix and its own measurement.
The next two chapters attack the two biggest: Chapter 8 on chunking, which determines what can ever be retrieved, and Chapter 9 on the two-stage retrieval that closes gaps 1 and 2 together.
RECALL
- Name the three retrieval gaps and one fix for each.
- A user’s question returns a wrong answer. What four checks tell you which gap you are in?
- Why is “fine-tune for form, retrieve for fact” the rule? Give a case where retrieval is required for legal rather than quality reasons.
- Which gap does reranking address, and which does it not?
- Why can chunk-based RAG not answer “what themes recur across these 200 reports?” — and what can?
- A team’s answer to poor retrieval is to upgrade the generation model. Why is this usually the wrong move?