In Java, a citation is a reference you trust because the compiler already
checked the types line up. An LLM's [2] is not that — it is a token the model
generated because a citation-shaped answer scored well, and nothing forces the sentence in
front of it to be true of chunk 2. Evaluating RAG scored
faithfulness for a whole answer at once; this exercise is the per-claim version that
actually tells you which sentence to fix. No judge model here — that comes next, in
ex-ragas-mock — because a deterministic check is the thing you calibrate a
judge model against, and a red row you can read beats a score you have to trust.
support(claim: str, chunk: str) -> float — a score in [0, 1]:
the fraction of the claim's content words (stopwords excluded) that also appear in the
chunk, multiplied down hard when the claim states a number the chunk does not contain. No
embeddings, no model — a reviewer reading the red row can see exactly which words or which
number failed to match.check(answer_with_cites: str, chunks: list[str]) -> Report — split the
answer into claims, each ending in a [n] marker exactly as it would read on a
page (1-indexed), and label each one "supported", "unsupported"
or "wrong-number" against the chunk its own marker names.The tests are ordinary pytest and ship in the public folder with the
starter — read them first; the names below are the check list. Solutions are not published.
This runs locally rather than in the browser because it is the production shape:
f1-build reuses this exact module to check citations on 50 real answers, and a
pytest suite with its own pyproject.toml is what that module actually ships as.
Needs git. uv installs the right Python itself, so nothing else is required.
# once, anywhere on your machine
git clone https://github.com/theDocWho/ai-ml-roadmap.git
cd ai-ml-roadmap
No git? Download the ZIP, unzip it, and cd into the unzipped folder instead.
From the repo root:
# one-time: uv (https://docs.astral.sh/uv/) manages the venv and pins Python ≥ 3.12 cd exercises/ex-citation-check && uv sync && uv run pytest -q # the same bar the reference solution clears uv run ruff check . && uv run mypy src
Done when uv run pytest -q prints 7 passed. Rerun
after every edit; pytest's -q output is the only readout this exercise has.
test_support_high_overlap_scores_full — a claim whose every content word
and number is in the chunk scores 1.0.test_support_disjoint_claim_scores_zero — a claim about a different topic
entirely shares no content word with the chunk, and scores 0.0.test_support_number_mismatch_is_penalized — a claim that shares almost every
word with its chunk but states the wrong number still scores well below the word-overlap
alone.test_support_stopwords_alone_score_zero — a claim built entirely of filler
words ("it", "was", "of", "that"…) scores 0.0 against any chunk, even one that happens to
share those same filler words.test_check_labels_supported_unsupported_wrong_number — three claims, three
chunks, and all three statuses in one report.test_check_citation_index_is_one_indexed — [2] names the
second chunk regardless of the order the claims are written in.test_check_out_of_range_citation_raises — a marker with no matching chunk
([0], or past the end of the list) raises instead of silently reading garbage.This is self-attestation — the site cannot see your terminal, so the box and the button are you telling The Path the suite went green on your machine.
\d+ is enough); if the claim's numbers are not a subset of
the chunk's, multiply the overlap score down — don't zero it, a claim can still be mostly
right with one wrong figure.[n] a claim ends in is written 1-indexed, the
way a citation reads on a page; the chunks list you index into is not. Write
the chunks[n - 1] down before you write anything else in check.wrong-number vs unsupported — decide status from the
word overlap first (below threshold is always "unsupported",
regardless of numbers), then split the rest by whether the claim's numbers are backed.uv run pytest -q -x --tb=short stops at the first
failure and shows the assertion that tripped; the message names the observed score or
status, not the fix.