Evaluating RAG

"Our RAG system is 68% accurate" is a number you cannot act on. A RAG system is two systems — a retriever that finds evidence and a generator that uses it — and a single end-to-end score is their product. Two very different failures produce the identical number, and the fix for one is worthless against the other. That is the entire argument for component metrics: measure the halves separately, because you can only fix them separately.

context recallfaithfulness answer relevanceRAGASgolden set

The same score from two different diseases

Below, 2000 questions run through a simulated RAG system with two independent dials: how often the retriever puts the right passage in the context (context recall), and how often the generator actually grounds its answer in what it was given rather than answering from memory (faithfulness). End-to-end accuracy is measured from the outcomes, not computed:

Two dials, one score

does the retriever put the evidence in the window?
does the generator use the evidence it was given?

Drag the two dials in opposite directions and watch the end-to-end number sit still. That is the whole problem with a single score: it is a curve of equally-plausible explanations, and picking the wrong point on that curve means a month spent tuning prompts when the retriever was never returning the document.

Which fix is worth doing?

The follow-up question is the one that matters on a Monday morning: given where you are, is your next ten points better spent on retrieval or on generation? The panel below measures the end-to-end gain from adding the same 10 points to each dial, from the current position:

+10 points of recall, or +10 of faithfulness?

The two levers are not symmetric, and that is the useful part. Faithfulness only pays off on questions where the evidence was there — its value is multiplied by recall — while recall returns something no matter how faithful the generator is. Put the two dials at 45/95 and retrieval returns more than three times what grounding does; move to 90/55 and grounding takes the lead. So the order of operations is not a matter of taste: get recall high enough that most questions have their evidence, then work on grounding. Doing it the other way round feels productive and moves very little.

What to actually measure, and on what

The metrics RAGAS and similar frameworks compute are these two halves plus a couple of neighbours, and each answers a specific question:

All of these need a golden set: questions with known answers and known supporting passages. Twenty to fifty real questions, written by hand from your own corpus, is enough to be useful and is the single highest-value hour in a RAG project. Without it you are comparing screenshots. And note the recursion — most of these metrics are computed by an LLM judge, which carries its own noise floor and biases; see LLM evaluation before trusting a two-point difference.

⚠️ Traps & honesty: this simulates a RAG system with two independent dials — real retrieval and generation quality are correlated, because a context full of near-misses actively misleads a generator rather than merely failing to help it · "faithfulness" here is modelled as a single probability of grounding, whereas real faithfulness is per-claim and partial · the model includes a small chance the generator answers correctly from parametric knowledge with no evidence at all, which is why the score never falls to zero — and which in real evaluation is a genuine confound, because a right answer for the wrong reason still scores · abstention is modelled as never wrong and never right; a system that abstains well is much better than these numbers suggest · sampling noise on 2000 questions is around a point, so treat small differences as noise.
Takeaways: a RAG system is two systems, and one end-to-end number is their product — the same score comes from a bad retriever with a faithful generator or the reverse · measure the halves separately: context recall and precision for retrieval, faithfulness and answer relevance for generation · the value of faithfulness is multiplied by recall, so when retrieval is weak, prompt work pays almost nothing — fix retrieval first · context precision matters as much as recall, because the right passage at rank 40 is nearly a miss · build a golden set of 20–50 hand-written questions with known supporting passages; it is the highest-value hour in the project · the metrics themselves are usually LLM-judged, so they carry a noise floor of their own. Next: advanced RAG.

Second opinion (taught here — these corroborate): RAGAS · Anthropic · contextual retrieval · Lost in the Middle.