Five topics on prompting, embeddings, vector search and agents. A RAG system is mostly retrieval, and retrieval is mostly the four decisions below. Get them wrong and the model hallucinates confidently over irrelevant context — which is the failure everyone ships and nobody notices until a user does.
You're building the layer that sits between a document store and the model. It has to cut documents into retrievable pieces, find the relevant ones, avoid handing the model five copies of the same paragraph, fit inside a context budget, and — the part most systems skip — say when it doesn't know.
chunk_text — fixed windows with overlap, so a sentence split across a
boundary still survives in one piece somewhere.cosine_topk — ranking by cosine, not by raw dot product. With unnormalised
vectors those are different questions, and one of them is "which document is longest".mmr_select — relevance and diversity. Pure top-k on a corpus with
near-duplicates returns the same fact five times and wastes the whole context window.build_prompt — fit the best chunks into a character budget without ever
exceeding it.should_answer — the refusal gate. If nothing retrieved is similar enough,
the honest output is "I don't know".The grader runs your code, then runs 6 checks against it with data you can't see — so solving the example instead of the problem will fail. Each check reports exactly what it expected and what it got. All 6 green marks this phase ready ✓ on your roadmap.
That you can build retrieval that is honest about its own limits. The three failures here — ranking by magnitude, filling the window with duplicates, and answering when nothing relevant was found — are the three reasons production RAG disappoints, and none of them is a model problem. They are all decisions in the layer you just wrote. The refusal gate in particular is what separates a system people trust from one they learn to double-check.