RAG Handbook retrieval-augmented generation Bipin Singh
Evaluation & production

Evaluation

2 min readChapter 18 of 26By Bipin Singh

RAG has many moving parts, and "it seems better" is not a strategy. Evaluation gives you numbers that tell you whether a change to chunking, retrieval, or prompting actually helped. Teams that measure improve steadily; teams that eyeball go in circles.

Evaluate the two halves separately

Because retrieval quality caps generation quality, measure them independently. If answers are bad, you need to know whether the right context wasn't retrieved (a retrieval problem) or the model mishandled good context (a generation problem). They have completely different fixes.

Retrieval metrics

MetricQuestion it answers
Recall@kOf the relevant chunks, how many made it into the top k?
Precision@kOf the top k retrieved, how many are actually relevant?
Hit rateDid at least one relevant chunk appear in the top k?
MRRHow high up was the first relevant chunk (mean reciprocal rank)?
nDCGAre the most relevant chunks ranked highest (order-aware)?

Recall is usually the one to watch first: if the answer isn't retrieved at all, nothing downstream can save you.

Generation metrics

Key idea
Faithfulness and answer relevancy are the two generation metrics to start with. The first catches made-up claims; the second catches fluent answers that dodge the question.

Build a golden dataset

Assemble a set of representative questions with known-good answers and the source passages that support them. This is your test suite. You can bootstrap it by having an LLM generate question–answer pairs from your documents, then curate. A few dozen good examples beat none; grow it over time, especially with real failures you find in production.

LLM-as-a-judge

Many quality metrics (faithfulness, relevancy) are graded by prompting a capable LLM to score the answer against the context and question. It's scalable and correlates reasonably with human judgement — but the judge has biases (it can favour longer or more confident answers), so calibrate it against human ratings on a sample and don't treat its scores as ground truth.

// Sketch of an LLM faithfulness check.
const verdict = await judge.complete(
  "Is every claim in the ANSWER supported by the CONTEXT? " +
  "Reply with a score 0-1 and list any unsupported claims.\n\n" +
  "CONTEXT:\n" + context + "\n\nANSWER:\n" + answer
);

Frameworks

You don't have to build all this by hand. Tools like RAGAS, TruLens, DeepEval, and Phoenix implement the standard metrics and integrate with pipelines and CI. Wire evaluation into your workflow so every change is scored automatically — evaluation you have to remember to run is evaluation you won't run.

The improvement loop

1. Build a golden set of questions + expected answers/sources.
2. Measure retrieval (recall@k) and generation (faithfulness, relevancy).
3. Change ONE thing (chunk size, hybrid, reranker, prompt).
4. Re-measure. Keep it if the numbers improve; revert if not.
5. Feed real production failures back into the golden set.
Tip
Change one variable at a time. If you swap the embedding model, chunking, and prompt at once and the score moves, you've learned nothing about which change did it.
Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me