Exercise ex-bm25 — BM25 from scratch, and how retrieval gets graded

In Java you'd reach for a search library and trust its ranking; here you build the ranking function yourself, because "the retriever found the wrong chunk" is the single most common way a RAG pipeline fails, and you can't fix what you can't compute by hand. Vector search ranked chunks by embedding distance; BM25 ranks them by term statistics — no model, no vectors, just word counts and how rare each word is. Most production retrievers run both and fuse the results, which is next.

~90 minruns in the browser 8 checksex-bm25

What you're building

Check 2 recomputes the exact BM25 formula independently and compares to your score() to 1e-6 — matching it means your idf, your tf-saturation term and your length-normalisation term are all individually correct, not just "close on this one example". Check 3 sets b=0 on two documents that share a term's frequency but differ only in padding: if your length ratio is right, the padding must not move the score at all.

If you get stuck