The Retrieval Gap

Lesson 07 / 16 · updated 2026-10-01 · 11 min


ONELINE

RAG fails in exactly three places, and naming which one you are in tells you which fix to reach for. Reaching for the wrong one is the most commonly wasted quarter in this field.

METAPHOR

Three ways to fail to find a book: you used the wrong words, it was on the wrong shelf, or you had it in your hands and skimmed past the page.

The idea behind retrieval-augmented generation is almost embarrassingly simple. The model does not know your private data. So look the relevant part up, put it in the prompt, and ask the question.

Simple to state, and it demos beautifully. It also fails in production constantly, and teams respond by swapping the model — which almost never helps, because the model is rarely the problem.

Why look things up instead of training it in?

The alternative is fine-tuning: bake the knowledge into the weights. The comparison is worth memorising:

Fine-tuning Retrieval
Where knowledge lives In the weights In the context
Updating it Retrain Update a row
Attribution None — a black box Explicit citations
Removing a document Very hard Delete it
Access control Nearly impossible per-user Filter at query time

The rule of thumb: fine-tune for form, retrieve for fact. Fine-tuning teaches style, tone, format, and task-shape. Retrieval supplies knowledge.

The rows about attribution and deletion are not soft benefits. In a regulated setting, “show me the source” and “delete this customer’s data” are legal requirements, and only one column can satisfy them.

How vector search actually works

“Look the relevant part up” is carrying a lot of weight in that sentence. Here is the architecture almost every system uses. It is called a bi-encoder, because the query and the documents pass through the encoder separately.

Bi encoderFlowchart with 10 labelled stages: Indexing time (once, offline; Document; Encoder; Vector stored in index; Query time (every request; Query; Same encoder; Query vector; Nearest-neighbour search; Ranked documents. Connections: Vector stored in index leads to Nearest-neighbour search; Query vector leads to Nearest-neighbour search; Nearest-neighbour search leads to Ranked documents.Query time (every request)Indexing time (once, offline)DocumentEncoderVector stored in indexQuerySame encoderQuery vectorNearest-neighbour searchRanked documents
The bi-encoder. Documents are encoded once, offline; the query is encoded at request time; comparison is pure geometry.

Documents are encoded ahead of time and stored. At query time you encode only the query, then find the nearest stored vectors. Chapter 2’s geometry is what makes “nearest” mean anything.

This shape dominates because it moves nearly all the work off the hot path. A million documents can be encoded overnight; a query needs one encoder pass and a nearest-neighbour lookup.

But look closely at what was given up. The document was encoded without ever seeing the query. The encoder had to guess, in advance, every question that document might answer, and compress that into one point.

Hold on to that sentence. It is the root of the first two gaps below, and the reason Chapter 9 puts a second, slower stage behind this one.

One choice is worth making deliberately. A purpose-built embedding model beats a language model you pooled yourself, essentially always. Embedding models are trained with a retrieval objective — pull matching query and document pairs together, push mismatched ones apart. A general-purpose model was never asked to make that arrangement true.

Measuring closeness

“Nearest” needs a definition. Three metrics show up, and one of them matters.

Metric Measures Use when
Cosine similarity Angle between vectors Default for text. Ignores length.
Dot product Angle and magnitude Vectors already normalised, or length is meaningful
Euclidean distance Straight-line gap Rare for text; common for images

Cosine is the default because document length should not make a document more or less relevant, and raw magnitude tends to track length. Cosine throws magnitude away and keeps direction, which is where the meaning is.

NAPKIN — Why normalised vectors make dot product free

For vectors already scaled to length 1, cosine similarity is the dot product — the denominator is 1 x 1.

So most vector databases normalise on write, then use dot product on read. Same ranking, one fewer square root per comparison, across billions of comparisons.

If you ever see a system storing normalised vectors but computing full cosine, you have found free performance.

The three gaps

When a RAG system gives a bad answer, exactly one of three things went wrong. Diagnosing which is the entire skill.

Three gapsFlowchart with 8 labelled stages: User question; Wording match the document?; GAP 1 - semantic mismatch Fix: hybrid search, query rewriting; Ranked into the top K?; GAP 2 - ranking failure Fix: rerank, retrieve more; Used in the answer?; GAP 3 - lost in the middle Fix: reorder, compress, fewer chunks; Good answer. Connections: Wording match the document? leads, No, to GAP 1 - semantic mismatch Fix: hybrid search, query rewriting; Wording match the document? leads, Yes, to Ranked into the top K?; Ranked into the top K? leads, No, to GAP 2 - ranking failure Fix: rerank, retrieve more; Ranked into the top K? leads, Yes, to Used in the answer?; Used in the answer? leads, No, to GAP 3 - lost in the middle Fix: reorder, compress, fewer chunks; Used in the answer? leads, Yes, to Good answer.NoYesNoYesNoYesUser questionWording matchthe document?GAP 1 - semantic mismatchFix: hybrid search, query rewritingRanked intothe top K?GAP 2 - ranking failureFix: rerank, retrieve moreUsed inthe answer?GAP 3 - lost in the middleFix: reorder, compress, fewer chunksGood answer
The three gaps. Each has a different fix, and each requires a different measurement to detect.

Gap 1: Semantic mismatch

The answer is in your corpus, but the query and the document do not look alike to the retriever.

The user asks about “fast cars”. The document says “the Porsche 911 accelerates from 0-60 in 3.5 seconds” and never uses the word fast. Or — more commonly in technical corpora — the user searches NVIDIA_VISIBLE_DEVICES and dense retrieval, which works on meaning rather than exact strings, shrugs at an identifier it has no semantic intuition for.

Fixes: hybrid search, which adds keyword matching alongside semantics (Chapter 9); query rewriting, where the model rephrases or expands the question before searching.

Gap 2: Ranking failure

The document was retrieved. It came back at position 47. You passed the top 10 to the model.

This is a precision problem, not a recall problem, and it is the one most often misdiagnosed — teams see “the right answer wasn’t in the prompt” and conclude the search is broken, when the search actually found it.

Fixes: reranking, which re-scores candidates with a slower, more accurate model (Chapter 9); retrieving more candidates before filtering.

Gap 3: Lost in the middle

The right chunk was in the prompt, in the top five. The model still ignored it.

This is Chapter 6’s problem: attention spread thin across a long context, relevant material buried in the middle, drowned out by nine other chunks of plausible-looking noise.

Fixes: send fewer, better chunks; reorder so the strongest land first and last; compress chunks before inserting them.

WATCHOUT

You cannot fix a gap you have not identified, and all three present identically to the user as “the answer was wrong.”

Instrument the pipeline so you can tell them apart. For a sample of failures, check: was the right document in the corpus at all? Did it appear anywhere in the retrieved candidates? What rank? Was it in the final prompt?

Those four checks split every failure into gap 1, gap 2, gap 3, or “we never had the document” — and each has a different owner and a different fix. Chapter 13 turns this into a standing measurement.

The four shapes of RAG

Systems are usually described by how much agency the retrieval step has. The progression is worth knowing because it is a ladder, and most teams should climb it slowly.

Naive RAG. Query → vector search → top-k → generate. This is the demo. It is also the version that produces all three gaps at once, and it should not be the target architecture for anything in production.

Advanced RAG. Query rewriting → hybrid search → rerank → generate. Multi-stage, still a straight line. This closes gaps 1 and 2 and is the sensible default for most production systems. Chapters 8 and 9 are mostly about building this well.

Agentic RAG. The model decides what to search for, evaluates what came back, and searches again if it is not satisfied. A loop rather than a pipeline. Far more capable on hard questions, and considerably more expensive and harder to predict — every extra loop is another full request. Part III is about loops.

GraphRAG. Extract entities and relationships into a knowledge graph, then traverse it. The distinctive win is aggregative questions — “what themes recur across these 200 incident reports?” — which chunk-based retrieval fundamentally cannot answer, because no single chunk contains the answer. The cost is a substantial extraction pipeline.

ASIDE

Do not start at agentic or graph. Build the advanced pipeline, measure your three gaps, and let the measurements tell you whether you need more machinery.

Choosing between them

Your situation Start here
Corpus under ~50k tokens, fairly stable No retrieval — cached long context (Chapter 6)
Standard question answering over documents Advanced RAG: hybrid + rerank
Questions needing several lookups to answer Agentic RAG
“Summarise across all documents” questions GraphRAG
Corpus over 100k tokens and changing hourly Advanced RAG, no question

What to carry forward

RAG is not one thing that works or does not. It is a pipeline with three identified failure points, each with its own fix and its own measurement.

The next two chapters attack the two biggest: Chapter 8 on chunking, which determines what can ever be retrieved, and Chapter 9 on the two-stage retrieval that closes gaps 1 and 2 together.

RECALL

  1. Name the three retrieval gaps and one fix for each.
  2. A user’s question returns a wrong answer. What four checks tell you which gap you are in?
  3. Why is “fine-tune for form, retrieve for fact” the rule? Give a case where retrieval is required for legal rather than quality reasons.
  4. Which gap does reranking address, and which does it not?
  5. Why can chunk-based RAG not answer “what themes recur across these 200 reports?” — and what can?
  6. A team’s answer to poor retrieval is to upgrade the generation model. Why is this usually the wrong move?