"Our RAG system is 68% accurate" is a number you cannot act on. A RAG system is two systems — a retriever that finds evidence and a generator that uses it — and a single end-to-end score is their product. Two very different failures produce the identical number, and the fix for one is worthless against the other. That is the entire argument for component metrics: measure the halves separately, because you can only fix them separately.
Below, 2000 questions run through a simulated RAG system with two independent dials: how often the retriever puts the right passage in the context (context recall), and how often the generator actually grounds its answer in what it was given rather than answering from memory (faithfulness). End-to-end accuracy is measured from the outcomes, not computed:
Drag the two dials in opposite directions and watch the end-to-end number sit still. That is the whole problem with a single score: it is a curve of equally-plausible explanations, and picking the wrong point on that curve means a month spent tuning prompts when the retriever was never returning the document.
The follow-up question is the one that matters on a Monday morning: given where you are, is your next ten points better spent on retrieval or on generation? The panel below measures the end-to-end gain from adding the same 10 points to each dial, from the current position:
The two levers are not symmetric, and that is the useful part. Faithfulness only pays off on questions where the evidence was there — its value is multiplied by recall — while recall returns something no matter how faithful the generator is. Put the two dials at 45/95 and retrieval returns more than three times what grounding does; move to 90/55 and grounding takes the lead. So the order of operations is not a matter of taste: get recall high enough that most questions have their evidence, then work on grounding. Doing it the other way round feels productive and moves very little.
The metrics RAGAS and similar frameworks compute are these two halves plus a couple of neighbours, and each answers a specific question:
All of these need a golden set: questions with known answers and known supporting passages. Twenty to fifty real questions, written by hand from your own corpus, is enough to be useful and is the single highest-value hour in a RAG project. Without it you are comparing screenshots. And note the recursion — most of these metrics are computed by an LLM judge, which carries its own noise floor and biases; see LLM evaluation before trusting a two-point difference.
Second opinion (taught here — these corroborate): RAGAS · Anthropic · contextual retrieval · Lost in the Middle.