Boss challenge: the retrieval layer that refuses to guess

Five topics on prompting, embeddings, vector search and agents. A RAG system is mostly retrieval, and retrieval is mostly the four decisions below. Get them wrong and the model hallucinates confidently over irrelevant context — which is the failure everyone ships and nobody notices until a user does.

chunking cosine retrieval MMR diversity context budget refusal

The situation

You're building the layer that sits between a document store and the model. It has to cut documents into retrievable pieces, find the relevant ones, avoid handing the model five copies of the same paragraph, fit inside a context budget, and — the part most systems skip — say when it doesn't know.

Dot product is not cosine unless you normalise. An embedding with a large magnitude wins every dot-product comparison regardless of direction — so the "most relevant" document becomes whichever one happens to have the biggest vector. Divide by the norms and you are comparing angle, which is what semantic similarity actually means.
Top-k on near-duplicates returns one fact, k times. Real corpora are full of repeated boilerplate. Ranked purely by relevance, the top 3 can be three versions of the same paragraph — the model sees one idea and none of the context that would have answered the question. MMR trades a little relevance for coverage, and it is usually a bargain.
A retrieval system that always answers is a hallucination generator. If the best match is barely similar to the query, returning it anyway hands the model irrelevant context and an implicit instruction to use it. A threshold and an honest refusal beat a confident invention every time.

Write it

The grader runs your code, then runs 6 checks against it with data you can't see — so solving the example instead of the problem will fail. Each check reports exactly what it expected and what it got. All 6 green marks this phase ready ✓ on your roadmap.

Phase 6 · boss challenge

Where is Python coming from? No server is involved. The browser downloads CPython compiled to WebAssembly the first time you press Run & grade, along with numpy, then runs your code locally. Nothing you write leaves the machine — which also means the grader is honest: it really executed what you wrote.

What passing actually proves

That you can build retrieval that is honest about its own limits. The three failures here — ranking by magnitude, filling the window with duplicates, and answering when nothing relevant was found — are the three reasons production RAG disappoints, and none of them is a model problem. They are all decisions in the layer you just wrote. The refusal gate in particular is what separates a system people trust from one they learn to double-check.

Take away: in RAG, the model is rarely the weak link. What you put in the context window decides the answer — including the decision not to answer at all.
Next: Phase 6B — inside an LLM