Evaluation: The Eval Is the Product

Lesson 13 / 16 · updated 2026-10-01 · 9 min


ONELINE

You cannot improve a system whose output you cannot score. In this field the eval harness, not the prompt, is the durable asset.

METAPHOR

A thermometer you had to build yourself, and now have to trust.

In ordinary software, correctness is mostly decidable. The function returns 4 or it does not. Your test suite tells you whether you broke something.

Here, the output is open-ended text, there are many acceptable answers, the same input can produce different outputs, and “better” is often a judgement call. Teams respond by not measuring — they look at a handful of outputs, decide it seems fine, and ship.

Then they cannot tell whether last week’s prompt change helped or hurt. That is the actual failure. Not the absence of tests; the absence of a ruler.

ASIDE

Models, frameworks, and vendors turn over fast. A well-built eval set for your domain stays valuable through all of it. It is the asset worth investing in.

Start with the dataset, not the metric

Before any tooling, you need examples: inputs paired with what good looks like.

A hundred well-chosen cases beat ten thousand scraped ones. Build it from:

  • Real user queries, especially the ones that went wrong. Production logs are your best source of realistic inputs.
  • Known failure modes. Every bug you fix becomes a permanent test case.
  • Edge cases — empty input, adversarial input, out-of-scope questions, ambiguous questions, questions your system should refuse.

This is unglamorous, mostly manual, and the highest-value work in the project. Budget real time for it. The teams that do this are the ones who can move fast later, because they can tell whether a change helped.

Four ways to score

Deterministic checks. Does it parse as JSON? Does it match the schema? Is it under the length limit? Does it contain the required citation? Cheap, fast, unambiguous. Use these for everything they can cover — which is more than people assume.

Reference-based metrics. Compare against a known-good answer. Works when there is genuinely one right answer — extraction, classification, structured output. Falls apart on open-ended generation, where a good answer worded differently scores badly.

LLM-as-judge. Use a model to score the output against criteria. The workhorse for open-ended quality, and the one that needs the most care.

Human review. The ground truth, and too slow and expensive to run continuously. Its real job is to calibrate the judge — you grade a sample by hand, check the judge agrees, and then trust the judge at scale.

Making LLM-as-judge trustworthy

A judge is just another model call, with all the same failure modes. The difference between a useful judge and a random number generator is discipline.

Score specific criteria, not vibes. Not “rate this 1–10”. Instead: is it factually correct given the context; does it address the question; does it cover every part; is it clear? Separate scores, each with a stated reason.

Force a justification before the score. Requiring the judge to explain first measurably improves consistency — and gives you something to read when you suspect it is wrong.

Use a coarse scale. Judges are unreliable at distinguishing a 7 from an 8. Binary pass/fail, or a 1–5 scale with explicit descriptions of what each point means, is far more stable.

Prefer comparison to absolute scoring. “Which of these two answers is better?” is a much easier question than “score this out of 10”, and judges are substantially more reliable at it.

WATCHOUT

Judges have systematic biases. They prefer longer answers. They prefer text in their own style. They favour the first option presented in a pairwise comparison — so randomise the order.

And a model judging its own family’s output tends to score it generously.

None of this makes judges useless. It makes them instruments that require calibration: grade 50 examples by hand, measure agreement with your judge, and fix the prompt until they mostly agree. An uncalibrated judge is a number that feels like measurement and is not.

Evaluating RAG: four metrics that fail independently

Chapter 7 said a RAG failure is really one of three gaps. Evaluation makes that diagnosis routine rather than heroic.

Rag metricsFlowchart with 7 labelled stages: Question; Retrieval; CONTEXT RECALL of what was needed, how much did we get?; CONTEXT PRECISION of what we retrieved, how much was relevant?; Generation; FAITHFULNESS is the answer supported by the context?; ANSWER RELEVANCY does it address the question asked?. Connections: Retrieval leads to CONTEXT RECALL of what was needed, how much did we get?; Retrieval leads to CONTEXT PRECISION of what we retrieved, how much was relevant?; CONTEXT RECALL of what was needed, how much did we get? leads to Generation; CONTEXT PRECISION of what we retrieved, how much was relevant? leads to Generation; Generation leads to FAITHFULNESS is the answer supported by the context?; Generation leads to ANSWER RELEVANCY does it address the question asked?.QuestionRetrievalCONTEXT RECALLof what was needed,how much did we get?CONTEXT PRECISIONof what we retrieved,how much was relevant?GenerationFAITHFULNESSis the answer supportedby the context?ANSWER RELEVANCYdoes it addressthe question asked?
Two retrieval metrics and two generation metrics. Which one drops tells you where to look.
Metric Question it answers What a low score means
Context recall Did we retrieve everything needed? Gap 1 — chunking or search is missing it
Context precision Was what we retrieved relevant? Gap 2 — ranking is poor, add reranking
Faithfulness Is the answer grounded in the context? The model is hallucinating beyond its sources
Answer relevancy Does it address the actual question? The model is answering a different question

The power here is separation. A system with high context recall and low faithfulness has a generation problem — the right information was there and the model invented something anyway. A system with high faithfulness and low context recall has a retrieval problem — the model is being faithful to material that does not contain the answer.

These require completely different fixes, and without separate metrics they look identical from the outside.

Offline and online

Offline evaluation runs your dataset through the system on every change. This is the gate: it runs in CI, and a regression on your core cases blocks the deploy. It is fast and repeatable and it does not reflect reality perfectly.

Online evaluation measures production: thumbs up and down, task completion, escalation to a human, whether the user rephrased immediately (a reliable signal that the first answer failed), and sampled judge scores on live traffic.

You need both. Offline catches regressions before users do. Online catches the gap between your dataset and the world — and feeds new cases back into the dataset, which is how the eval set stays honest.

The score tells you that; the trace tells you where

A score is a verdict. On its own it is not actionable: faithfulness fell from 0.86 to 0.71 and you have no idea which of a dozen moving parts moved.

The fix is to record the whole chain for every request, not a summary of it:

request id, tenant, feature
  model, prompt version, temperature
  original query -> rewritten query
  retrieved chunk ids, with retrieval and rerank scores
  the final assembled prompt
  each tool call, its arguments and its result
  raw output, before guardrails
  latency and tokens per stage
  cost, judge scores, final outcome

That record is an AI trace, and it is the join between the four things this part of the book treats separately. The eval says quality dropped. The trace says the reranker started timing out, so stage 2 was silently skipped, so gap 2 reopened. Chapter 15’s cost attribution reads the same rows. Chapter 14’s incident review reads them too.

Two properties are worth insisting on.

Sample everything, keep the failures. Full traces on every request get expensive fast. Sample the successes; retain every low score, every refusal, every guardrail trip, every retry. Those are the rows you will actually open.

Version the prompt, not just the code. A prompt change is a deploy. If your trace cannot tell you which prompt produced a response, you cannot attribute a regression to it, and prompt changes are the most frequent change you will make.

WATCHOUT

Instrument before you need it, because you cannot trace the past.

The moment you need a trace is an incident — quality has dropped, or a bill has tripled, and the requests that would explain it were served last week. If they were not recorded then, the only remaining option is to reproduce a non-deterministic system from memory. Teams discover this exactly once.

What to carry forward

Build the dataset first. Use deterministic checks wherever they reach, and calibrate your judge against human grades before trusting it. Score retrieval and generation separately so failures point at their own cause. Gate deploys offline; learn from production online.

Next: what to do about the failures your evals find, and the ones they do not.

RECALL

  1. Why is “we looked at some outputs and they seemed fine” not evaluation?
  2. Name the four scoring approaches and what each is best at.
  3. Give three ways to make an LLM judge more reliable, and three biases to guard against.
  4. Context recall is high and faithfulness is low. What is broken, and what is not?
  5. Why do you need both offline and online evaluation?
  6. What should a trace contain so that a bad score is diagnosable?