Exercise ex-rrf-mmr — fusing rankings without comparable scores, then re-ranking for diversity

In Java, combining two ranked lists feels like a merge: sort by whichever Comparator, or average two numbers that both look like "how good". But a BM25 score and a cosine similarity are not on the same scale — one is unbounded and log-shaped, the other lives in [-1, 1] — averaging them is often meaningless. Reciprocal Rank Fusion (RRF) sidesteps the mismatch entirely: it never touches a raw score, only where each document landed in each list. ex-bm25's rank() output and a vector search's ranking fuse here exactly as they will in f1-build's hybrid retriever. Once you have one fused list, Maximal Marginal Relevance (MMR) re-ranks it again — trading a little relevance for diversity, so the top slots aren't five near-duplicate chunks of the same paragraph.

~75 minruns in the browser 7 checksex-rrf-mmr

What you're building

Check 1 is a genuine trap for the plausible-but-wrong "sum the raw ranks" implementation. Rank sums and reciprocal-rank sums both reward "ranked highly everywhere", so on many small examples they agree; they part ways when one document has a very good rank in one list and a very ordinary rank in the other. Here doc 0 lands 1st and 6th while doc 1 lands 3rd and 3rd: a rank sum prefers doc 1 (6 vs 7), reciprocal rank prefers doc 0 (0.643 vs 0.5), because 1/(k+rank) rewards a top rank far more than it punishes a middling one. Check 3 is the same idea from a different angle — push k from 1 to 1000 and watch which document wins flip.

If you get stuck