Advanced RAG

Once you have measured that retrieval is the weak half, there are two standard moves, and both work by spending compute you were not spending. Re-ranking admits that the cheap scorer is noisy and pays a better one to reorder its shortlist. Query rewriting admits that the user's words are not the document's words and asks the question several ways. Both are worth their cost, and both have a ceiling worth knowing about before you build them.

cross-encoderre-ranking multi-queryHyDEhybrid search

Retrieve wide, re-rank narrow

A bi-encoder — the ordinary vector search — embeds the query and every document separately, so it can score a million documents in milliseconds and is correspondingly rough. A cross-encoder reads the query and one document together and scores the pair, which is far more accurate and far too slow to run over a corpus. The standard arrangement uses each for what it is good at: the bi-encoder proposes K candidates, the cross-encoder reorders them, and you keep the top 5. Below, 1200 queries against a 2000-document corpus — three documents genuinely answer each query, and forty are near-misses designed to fool a cheap scorer:

How wide should the shortlist be?

K — retrieve this many with the cheap scorer, keep the best 5 after re-ranking

The dashed ceiling is the thing to internalise: a re-ranker can only reorder what it was given. If the right document is not in the bi-encoder's top K, no amount of cross-encoder quality will find it — which is why the curve tracks the ceiling and then flattens against it. Widening K raises the ceiling and costs a cross-encoder call per candidate, so the shape of the decision is: push K up until the ceiling stops moving, then stop, because everything past that is paying for reordering documents that were already going to lose.

The other half of the argument is about where in the list the right chunk lands. Even when it is retrieved, a passage buried at rank 40 of 50 is much less useful than the same passage at rank 2 — attention over a long context is measurably non-uniform, which the Lost in the Middle work documented. So re-ranking improves the generator's input twice: it changes what is in the window, and it changes what is at the top of it.

When the user's words are not the document's words

The second failure has nothing to do with scoring quality. Some fraction of queries simply do not share vocabulary with the passage that answers them — the user says "how do I stop the charge going through", the document says "cancelling a pending transaction". Multi-query retrieval asks an LLM for several rephrasings, retrieves for each, and unions the results:

Asking the same question several ways

one LLM call to generate them, then one retrieval each
share of queries whose words do not match the answering passage

Diminishing returns arrive quickly, which is the practical point: three or four rephrasings capture most of what is available, and the curve is flat after that while the cost keeps rising linearly. It also does nothing for the queries that were already fine, so the whole benefit is concentrated in the mismatched slice — which means the honest way to size this feature is to measure how big that slice actually is on your corpus.

Three siblings of the same idea, in rough order of how often they earn their keep:

⚠️ Traps & honesty: both scorers are simulated as "true relevance plus Gaussian noise", with the cross-encoder's noise about a quarter of the bi-encoder's — real cross-encoders are better in a more structured way than a smaller sigma implies, and the ratio varies by model and domain · the corpus has exactly three genuinely relevant documents per query, which is tidier than reality · re-ranking latency is not modelled: a cross-encoder over 100 candidates is a real user-visible delay, and it is the main reason people cap K far below where the recall curve flattens · multi-query is modelled as independent draws at the mismatch step, so it looks more reliable than it is — real rephrasings are correlated, because they come from one model reading one question · none of these fix a chunking strategy that split the answer in half.
Takeaways: a bi-encoder is fast and rough, a cross-encoder is accurate and slow — retrieve K with the first, reorder with the second, keep 5 · the re-ranker can only reorder what it was given, so the bi-encoder's recall@K is a hard ceiling and the curve flattens against it · widen K until the ceiling stops moving, then stop; each extra candidate is another cross-encoder call · re-ranking also puts the right passage near the top, which matters because attention over long context is not uniform · a few rephrasings recover most of the vocabulary-mismatch losses and the curve is flat after three or four · hybrid keyword+vector search is the cheapest large win, because the two channels fail on different queries. Next: stateful agents.

Second opinion (taught here — these corroborate): Anthropic · contextual retrieval · Sentence-Transformers · retrieve & re-rank · Lost in the Middle.