Observability & monitoring
When a RAG answer is wrong, the failure could be in any stage. Observability means you can open up a request and see exactly what happened at each step — and monitoring means you're alerted when quality slips over time instead of hearing it from an angry user.
Trace every stage
Log the full journey of each request so you can reconstruct any answer:
- The original query and any rewritten/expanded versions.
- Which chunks were retrieved, their scores, and their sources.
- Reranking scores, if used.
- The exact final prompt sent to the LLM.
- The generated answer, token counts, latency per stage, and cost.
Monitor in production
- Operational: latency (p50/p95/p99), error and timeout rates, cost per query, throughput.
- Quality: retrieval scores, rate of low-confidence/empty retrievals, "I don't know" frequency.
- Drift: as content and query patterns change, retrieval quality can quietly degrade — watch trends, not just snapshots.
Close the loop with user feedback
Add thumbs up/down and, where possible, capture why. Negative feedback is gold: it points to real failures you can turn into new evaluation cases and prioritise fixes against. A steady feedback → golden-set → fix loop is how RAG systems get better after launch, not just before it.
Continuous evaluation
Run your evaluation suite in CI on every change so regressions are caught before deploy, and sample live traffic for periodic offline scoring so you notice drift in the wild. Tools like LangSmith and Phoenix are built to trace, evaluate, and monitor RAG and agent pipelines end to end.
Alerting
Set thresholds and alert on breaches: latency spikes, cost surges, a jump in empty retrievals, or a drop in average retrieval score. The goal is to learn about degradation from your dashboards, not your users.