Evaluating RAG named faithfulness, answer
relevance and context precision as RAGAS's three headline metrics, each one backed by an LLM judge call.
ex-citation-check was the deterministic warm-up:
no judge, a red row you can read. This exercise is the real shape — three metrics computed
from a judge's answers — but the judge is still deterministic: MockJudge
answers from a fixed fixture instead of calling a model, so every score below is exact, not
"roughly right". That is also the calibration move a real eval harness needs: build the
arithmetic against a judge you can trust completely, then swap in the real one.
faithfulness(row: EvalRow, judge: MockJudge) -> float — the fraction of the
row's statements that the judge says some context chunk supports; an answer with no
statements at all has stated nothing unfaithful.answer_relevance(row: EvalRow, judge: MockJudge) -> float — the average
similarity between the row's question and each question the judge reckons the answer is
actually answering; zero generated questions scores 0.0, not a divide-by-zero.context_precision(row: EvalRow, judge: MockJudge) -> float — RAGAS's average
precision over the row's contexts in rank order: a relevant chunk ranked first scores
higher than the same chunk ranked last, and a docset with no relevant chunk scores 0.0.run(dataset: list[EvalRow], judge: MockJudge) -> str — a markdown table, one
row per dataset entry plus a final mean row, every score to 2 decimal places.MockJudge, EvalRow and the similarity helper
ship already written in the starter — read them first, they're the contract the four
functions above are graded against. The tests are ordinary pytest and ship in the public
folder with the starter; the names below are the check list. Solutions are not published.
This runs locally, same as ex-citation-check, because f1-build
reuses this exact module for its RAGAS report.
Needs git. uv installs the right Python itself, so nothing else is required.
# once, anywhere on your machine
git clone https://github.com/theDocWho/ai-ml-roadmap.git
cd ai-ml-roadmap
No git? Download the ZIP, unzip it, and cd into the unzipped folder instead.
From the repo root:
# one-time: uv (https://docs.astral.sh/uv/) manages the venv and pins Python ≥ 3.12 cd exercises/ex-ragas-mock && uv sync && uv run pytest -q # the same bar the reference solution clears uv run ruff check . && uv run mypy src
Done when uv run pytest -q prints 6 passed. Rerun
after every edit; pytest's -q output is the only readout this exercise has.
test_faithfulness_scores_fraction_supported — every statement supported
scores 1.0; one of three unsupported scores 2/3.test_faithfulness_empty_statements_scores_one — an answer with zero
statements is vacuously faithful.test_answer_relevance_close_paraphrase_scores_higher_than_unrelated_answer —
a close paraphrase scores well above an unrelated answer, and an answer absent from the
judge's fixture scores 0.0 rather than raising.test_context_precision_relevant_chunk_ranked_first_beats_ranked_last — the
same single relevant chunk scores 1.0 at rank 1 and 0.5 at rank 2 — the metric is
rank-sensitive, not just a relevant-count ratio.test_context_precision_no_relevant_chunk_scores_zero — no chunk relevant
scores 0.0 without dividing by zero.test_run_writes_markdown_table_with_mean_row — a two-row dataset produces
exact per-row scores and a correctly averaged **mean** row.This is self-attestation — the site cannot see your terminal, so the box and the button are you telling The Path the suite went green on your machine.
row.contexts[0]. Loop the statements outside, the contexts inside.judge.questions_from(row.answer) can return an
empty list; average over an empty list is the divide-by-zero this check exists to catch.f"{x:.2f}") before
adding the mean row; the mean row uses the exact same formatting, just averaged first.uv run pytest -q -x --tb=short stops at the first
failure and shows the assertion that tripped; the message names the observed score, not the
fix.