RAG Reranking: How to Improve Retrieval Accuracy in AI Applications
How reranking improves RAG: two-stage retrieval, cross-encoder and LLM rerankers, candidate set size, relevance scores and thresholds, latency and cost, and how to evaluate the gain.
Quick answer
Reranking adds a precise second stage to RAG retrieval. A fast first stage (vector, keyword or hybrid search) gathers a broad candidate set, perhaps the top 50 passages; a reranker, usually a cross-encoder that reads the query and each passage together, re-scores them; and only the top few go to the language model. This often fixes cases where the right passage was retrieved but ranked too low. Tune the candidate count against latency, consider score thresholds for refusals and measure the gain on your own question set.
Where This Fits
Reranking follows hybrid search in the RAG pipeline. Its gains depend on good chunking, since a reranker cannot fix a chunk that lacks the answer.
Why First-Stage Retrieval Is Not Enough
Embedding search compares a query vector with passage vectors computed independently, which is fast but approximate. Keyword search matches terms but not intent. Both produce reasonable candidate sets with imperfect order. Because the language model only sees the top few passages, a relevant passage at rank 15 might as well not exist. Reranking reorders the candidates using a model that reads query and passage together.
How Two-Stage Retrieval Works
- Stage 1, recall: hybrid search returns a candidate set (for example 30 to 100 passages) with permission filters applied
- Stage 2, precision: the reranker scores each candidate against the query
- Selection: keep the top few, optionally only those above a calibrated score threshold
- Generation: pass the selected passages to the model with instructions to cite them
Types of Rerankers
| Type | How it works | Fits |
|---|---|---|
| Cross-encoder models | Score query and passage pairs | Most RAG systems |
| Hosted rerank APIs | Managed cross-encoder style models | Teams avoiding model hosting |
| Late-interaction models | Token-level matching | Higher precision with some speed |
| LLM-based reranking | Language model judges or orders passages | Small candidate sets, offline work |
| Rules and signals | Boost recency, authority or source type | Combined with model scores |
Right documents retrieved but answers still wrong?
ZSpace Labs can add and tune reranking in your RAG pipeline and measure the improvement on your own questions.
Tuning Candidate Counts and Thresholds
The candidate set must contain the right passage for reranking to help, so measure first-stage recall at different sizes. Larger sets improve the chance but add latency and cost. After reranking, choose how many passages to pass on: too few risks missing context, too many adds noise and tokens. Calibrated score thresholds can trigger 'I could not find this in our documents' rather than a weak answer.
Combining Reranking With Business Signals
Relevance is not the only factor. A newer policy should outrank an older version; an official handbook should outrank a chat message. Combine reranker scores with metadata such as date, document status and source authority, or filter outdated documents before reranking.
Evaluating Reranking
Use the same question set as for retrieval evaluation. Compare, with and without reranking: the rank of the correct passage, recall at the number of passages you send to the model, answer faithfulness and correctness, and latency and cost per query. Keep reranking only if gains are clear on your data.
Advantages and Limitations
| Advantages | Limitations |
|---|---|
| Better ordering of retrieved passages | Adds latency and cost per query |
| Fewer, more relevant passages in the prompt | Cannot recover passages the first stage missed |
| Scores can support refusal thresholds | Thresholds need calibration |
| Easy to add to an existing pipeline | Another model to host or pay for |
How to Add Reranking Step by Step
- 1. Measure first-stage recall at several candidate sizes
- 2. Choose a reranker that fits latency, language and hosting needs
- 3. Rerank the candidate set and select the top passages
- 4. Compare metrics with and without reranking
- 5. Tune candidate counts and thresholds
- 6. Monitor latency and cost in production
Reranking, Filters and Permissions
Apply permission and metadata filters before reranking, in the first-stage retrieval. Reranking unfiltered candidates and filtering afterwards wastes compute and risks exposing restricted content in logs or debug views. If filters remove most candidates, increase the first-stage candidate count for filtered queries rather than reranking fewer results. Keep the filter logic in the retrieval service so every application using it inherits the same rules; see enterprise RAG architecture.
Tools and Hosting Options
Rerankers are available as hosted APIs from several model providers, as open-source cross-encoder models you can run on your own infrastructure, and built into some search engines and vector databases. Hosted APIs are fastest to adopt; self-hosting gives data control and predictable cost at volume but requires GPU or optimized CPU serving. Whichever you choose, check language support for your content and measure latency at your candidate set size.
Worked Example
An illustrative scenario, not a client case: a software company's documentation assistant retrieves the right troubleshooting page in its top 20 for most questions, but often not in the top 3 that reach the model. Adding a cross-encoder reranker over the top 40 candidates moves the correct page into the top 3 much more often in evaluation, with acceptable added latency.
Common Mistakes
- Adding reranking without measuring first-stage recall
- Reranking too few candidates to matter
- Passing many reranked passages anyway, adding noise
- Ignoring document freshness and authority
- Not tracking added latency
Want measurably better retrieval?
Talk to ZSpace Labs about RAG optimization and development.
Conclusion
Reranking separates finding candidates from ordering them, which often fixes the gap between 'retrieved somewhere' and 'used in the answer'. Measure recall, tune candidates and keep it only where it helps. Related: hybrid search and RAG guide.
Common questions
A second retrieval stage in which a more precise model re-scores the candidates returned by the first stage, so the most relevant passages are placed at the top and passed to the language model.