Evaluating RAG systems
How to measure RAG: retrieval quality, answer faithfulness, citation correctness, latency, and cost—not vibes.
What it is
How to measure RAG: retrieval quality, answer faithfulness, citation correctness, latency, and cost—not vibes.
Why it matters
RAG fails quietly. Without evals, prompt tweaks become superstition.
How it works (plain)
Build a golden set of questions with expected doc IDs and answer points. Track recall@k, groundedness, and human grades. Regress on every chunker/model change.
Everyday example
A open-book quiz graded on using the right pages—not only sounding smart.
Try it
Write 10 questions for your docs with the expected source filename for each.
Myths
- ⚠️ Myth: LLM-as-judge alone is enough.
- ✓ Reality: Useful with calibration; keep human spots and exact checks.
- ⚠️ Myth: ChatGPT-general benchmarks replace domain goldens.
- ✓ Reality: Your corpus needs your questions.
Sources
- Course 09 groundedness; Course 18 evaluation
- NIST AI RMF: https://www.nist.gov/itl/ai-risk-management-framework ↗
