Evaluation
RAG has many moving parts, and "it seems better" is not a strategy. Evaluation gives you numbers that tell you whether a change to chunking, retrieval, or prompting actually helped. Teams that measure improve steadily; teams that eyeball go in circles.
Evaluate the two halves separately
Because retrieval quality caps generation quality, measure them independently. If answers are bad, you need to know whether the right context wasn't retrieved (a retrieval problem) or the model mishandled good context (a generation problem). They have completely different fixes.
Retrieval metrics
| Metric | Question it answers |
|---|---|
| Recall@k | Of the relevant chunks, how many made it into the top k? |
| Precision@k | Of the top k retrieved, how many are actually relevant? |
| Hit rate | Did at least one relevant chunk appear in the top k? |
| MRR | How high up was the first relevant chunk (mean reciprocal rank)? |
| nDCG | Are the most relevant chunks ranked highest (order-aware)? |
Recall is usually the one to watch first: if the answer isn't retrieved at all, nothing downstream can save you.
Generation metrics
- Faithfulness / groundedness — is every claim in the answer supported by the retrieved context? Directly measures hallucination.
- Answer relevancy — does the answer actually address the question asked?
- Context precision — is the retrieved context on-topic, or padded with noise?
- Context recall — did retrieval capture everything needed to answer fully?
- Correctness — does the answer match a known ground-truth answer?
Build a golden dataset
Assemble a set of representative questions with known-good answers and the source passages that support them. This is your test suite. You can bootstrap it by having an LLM generate question–answer pairs from your documents, then curate. A few dozen good examples beat none; grow it over time, especially with real failures you find in production.
LLM-as-a-judge
Many quality metrics (faithfulness, relevancy) are graded by prompting a capable LLM to score the answer against the context and question. It's scalable and correlates reasonably with human judgement — but the judge has biases (it can favour longer or more confident answers), so calibrate it against human ratings on a sample and don't treat its scores as ground truth.
// Sketch of an LLM faithfulness check.
const verdict = await judge.complete(
"Is every claim in the ANSWER supported by the CONTEXT? " +
"Reply with a score 0-1 and list any unsupported claims.\n\n" +
"CONTEXT:\n" + context + "\n\nANSWER:\n" + answer
);
Frameworks
You don't have to build all this by hand. Tools like RAGAS, TruLens, DeepEval, and Phoenix implement the standard metrics and integrate with pipelines and CI. Wire evaluation into your workflow so every change is scored automatically — evaluation you have to remember to run is evaluation you won't run.
The improvement loop
1. Build a golden set of questions + expected answers/sources.
2. Measure retrieval (recall@k) and generation (faithfulness, relevancy).
3. Change ONE thing (chunk size, hybrid, reranker, prompt).
4. Re-measure. Keep it if the numbers improve; revert if not.
5. Feed real production failures back into the golden set.