Project — Flagship 1: a RAG + agent service with evals and tracing

Every exercise on this phase built one piece in isolation — BM25 ranking, RRF fusion, a citation checker, an MCP tool server, a RAGAS-style harness against a mock judge. This project wires them into one real service: ingest a document set into Postgres with pgvector, retrieve hybrid (BM25 + vectors, fused and re-ranked), answer through an agent loop that calls at least two tools, trace every call in Langfuse, and grade the whole thing with a RAGAS report and a citation check over 50 real answers. It runs on your machine and, later, on Fargate — this page is the checklist, not a tutorial to copy.

~12 hruns locally · python + docker 8 milestonesf1-build

What you're building

The milestones

Suggested build order

  1. Ingest first, and check it by hand — query a few chunks back before wiring anything else to it.
  2. Hybrid retrieval next, against the ingested corpus — this is where exercise-bm25 and exercise-rrf-mmr move from a browser exercise into service code.
  3. The agent loop and its two MCP tools, once retrieval returns real chunks to reason over.
  4. Langfuse tracing wrapped around the retrieval and tool calls you now have, so every later change is visible in a trace, not just believed.
  5. The evals last, against 50 real answers the running service produced — a RAGAS report and a citation check computed on fixtures would not prove anything about the service.
  6. Containerizing (docker compose: api + pgvector + Langfuse) and a Fargate deployment follow as their own cloud guides later on this track; this milestone list is satisfied by the service running locally first.

This is self-attestation — the site cannot see your repo or your Langfuse project, so the box and the button are you telling The Path the service exists and the milestones are real.

If you get stuck