Once you have measured that retrieval is the weak half, there are two standard moves, and both work by spending compute you were not spending. Re-ranking admits that the cheap scorer is noisy and pays a better one to reorder its shortlist. Query rewriting admits that the user's words are not the document's words and asks the question several ways. Both are worth their cost, and both have a ceiling worth knowing about before you build them.
A bi-encoder — the ordinary vector search — embeds the query and every document separately, so it can score a million documents in milliseconds and is correspondingly rough. A cross-encoder reads the query and one document together and scores the pair, which is far more accurate and far too slow to run over a corpus. The standard arrangement uses each for what it is good at: the bi-encoder proposes K candidates, the cross-encoder reorders them, and you keep the top 5. Below, 1200 queries against a 2000-document corpus — three documents genuinely answer each query, and forty are near-misses designed to fool a cheap scorer:
The dashed ceiling is the thing to internalise: a re-ranker can only reorder what it was given. If the right document is not in the bi-encoder's top K, no amount of cross-encoder quality will find it — which is why the curve tracks the ceiling and then flattens against it. Widening K raises the ceiling and costs a cross-encoder call per candidate, so the shape of the decision is: push K up until the ceiling stops moving, then stop, because everything past that is paying for reordering documents that were already going to lose.
The other half of the argument is about where in the list the right chunk lands. Even when it is retrieved, a passage buried at rank 40 of 50 is much less useful than the same passage at rank 2 — attention over a long context is measurably non-uniform, which the Lost in the Middle work documented. So re-ranking improves the generator's input twice: it changes what is in the window, and it changes what is at the top of it.
The second failure has nothing to do with scoring quality. Some fraction of queries simply do not share vocabulary with the passage that answers them — the user says "how do I stop the charge going through", the document says "cancelling a pending transaction". Multi-query retrieval asks an LLM for several rephrasings, retrieves for each, and unions the results:
Diminishing returns arrive quickly, which is the practical point: three or four rephrasings capture most of what is available, and the curve is flat after that while the cost keeps rising linearly. It also does nothing for the queries that were already fine, so the whole benefit is concentrated in the mismatched slice — which means the honest way to size this feature is to measure how big that slice actually is on your corpus.
Three siblings of the same idea, in rough order of how often they earn their keep:
Second opinion (taught here — these corroborate): Anthropic · contextual retrieval · Sentence-Transformers · retrieve & re-rank · Lost in the Middle.