# ex-ragas-mock — a RAGAS-style eval harness, judged by a fixture instead of a model

`ex-citation-check` caught one failure mode (a claim citing a chunk that doesn't back it up).
RAGAS's three headline metrics — faithfulness, answer relevance, context precision — catch the
rest, but a real judge model makes every score noisy. `MockJudge` here answers from a fixed
fixture instead of calling anything, so the numbers below are exact, not "roughly right".

## What to implement

`src/ragas_mock/core.py` (`MockJudge`, `EvalRow` and `_similarity` are provided):

- `faithfulness(row, judge) -> float` — the fraction of `row.statements` that `judge.supports`
  against some chunk in `row.contexts`.
- `answer_relevance(row, judge) -> float` — the average similarity between `row.question` and
  each question `judge.questions_from(row.answer)` reckons the answer is actually answering.
- `context_precision(row, judge) -> float` — RAGAS's average precision over `row.contexts` in
  rank order: a relevant chunk ranked first scores higher than the same chunk ranked last.
- `run(dataset, judge) -> str` — a markdown table, one row per dataset entry plus a `**mean**`
  row, scores to 2 decimal places.

## Run it

```
cd exercises/ex-ragas-mock
python3 -m venv /tmp/lbv-ragas && /tmp/lbv-ragas/bin/pip install -q pytest
/tmp/lbv-ragas/bin/python -m pytest -q
```

(`uv sync && uv run pytest -q` works too.)

## If you get stuck

- **`faithfulness`** — "supported by *some* chunk", not "supported by every chunk" and not
  "supported by `row.contexts[0]`". An empty `row.statements` is vacuously faithful: 1.0.
- **`answer_relevance`** — the judge's generated questions stand in for "what is this answer
  actually about"; a question with no generated paraphrases at all (an empty answer, or a fixture
  gap) scores 0.0 — average over zero questions is not "skip the row".
- **`context_precision`** — write out precision@k for `k = 1..len(contexts)` by hand for a
  3-chunk example with the relevant chunk in each position before writing the loop. The chunk's
  *rank* is the whole point of this metric; a version that only counts how many chunks are
  relevant, ignoring where they sit in the list, passes the average-of-relevant-chunks arithmetic
  but fails the ranking test here.
- **`run`** — build the row strings first and get those exact, then add the mean row last; both
  use the same `f"{x:.2f}"` formatting.
