Reranking in RAG
Reranking in RAG is a second retrieval stage that rescores the candidate passages returned by a fast first-stage search. A slower but more precise relevance model reorders these candidates, so the passage that most directly answers the question appears first in the prompt.
- Candidate retrieval: A fast retriever, such as vector search, returns a candidate list of 20 to 100 chunks.
- Relevance scoring: A reranker compares each candidate with the question and assigns a new relevance score.
- Cross-encoder: A transformer model that reads the question and a passage as one input and outputs a single relevance score.
- Top-k selection: Only the highest-scoring 3 to 5 passages after reranking are added to the prompt.
- Primary objective: Higher precision at the top of the ranked list, where the language model concentrates its attention.
For example, a developer asks what a client should do after a rate limit error, and the reranker moves the Rate limits section of the Orders API documentation above the Error codes section.
Key Characteristics of Reranking in RAG
- Two-stage retrieval: An inexpensive initial search provides recall, and the reranker adds precision on a short candidate list.
- Joint encoding: The reranker processes the question and passage together, so it can evaluate word order and exact phrases.
- Unmodified index: Reranking operates at query time and requires no modification to the stored embeddings.
- Interchangeable scorer: Any model that can rerank retrieved documents against the question can serve, such as a cross-encoder, a rule-based scoring function or an LLM prompt.
- Predictable latency: Processing time grows with the number of candidates, not with the size of the entire index.
How Reranking in RAG Works
- Candidate retrieval: The first stage queries the vector database, or a hybrid search index, and returns the top candidates.
- Pair construction: The system combines the question with each candidate passage to form question and passage pairs.
- Relevance scoring: The reranker assigns each pair a relevance score, usually a probability between 0 and 1.
- Sorting and selection: Candidates are sorted by the new score, and only the highest-ranked few are retained.
- Prompt assembly: The retained passages enter the prompt in ranked order, following the flow described in how RAG works.
Scoring and Cost of Two-Stage Reranking
The bi-encoder vs cross-encoder distinction is a difference in where the question and the passage meet. A bi-encoder encodes each text separately and compares the two vectors, while a cross-encoder reads both texts together and produces one relevance score:
- : the user's question, and : one candidate passage from the Orders API documentation.
- : the bi-encoder, an embedding model that maps any text to a vector independently of other texts.
- : cosine similarity, the angle-based comparison described in vector search.
- : the cross-encoder, a transformer that reads the question and the passage as one combined input.
In words, the bi-encoder allows every passage representation to be precomputed during indexing, so answering a query requires only one additional embedding and an inexpensive vector lookup. The cross-encoder must execute the complete transformer again for every question and passage combination, because its relevance score depends on both texts simultaneously. The per-query latency of each design is therefore:
- : the number of chunks in the index, and : the number of candidates passed to the reranker.
- : the latency of embedding the question, and : the latency of the approximate nearest-neighbour search.
- : the latency of one cross-encoder inference over a single question and passage pair.
The worked example uses illustrative timings for an index of 20,000 Orders API chunks:
- Assumed timings: Question embedding requires 10 ms, the vector search requires 5 ms, and each individual cross-encoder inference requires 4 ms.
- Two-stage retrieval: Reranking 50 candidates adds 200 ms to the 15 ms first stage, which gives a total of 215 ms per query.
- Cross-encoder only: Scoring every chunk with the cross-encoder requires 80,000 ms, which is entirely impractical for an interactive documentation assistant.
# Illustrative per-query costs in milliseconds for the Orders API assistant.
N = 20_000 # chunks in the index
EMBED_QUERY = 10.0 # one bi-encoder pass over the question
ANN_SEARCH = 5.0 # approximate nearest-neighbour lookup over all N chunks
CROSS_PAIR = 4.0 # one cross-encoder pass over a (question, chunk) pair
def two_stage(k):
return EMBED_QUERY + ANN_SEARCH + k * CROSS_PAIR
rows = [("bi-encoder only", EMBED_QUERY + ANN_SEARCH)]
rows += [(f"two-stage, k = {k}", two_stage(k)) for k in (20, 50, 100)]
rows += [("cross-encoder on all N", N * CROSS_PAIR)]
for label, ms in rows:
print(f"{label:<23} {ms:>7,.0f} ms")
print(f"speed-up at k = 50: {N * CROSS_PAIR / two_stage(50):,.0f}x")bi-encoder only 15 ms
two-stage, k = 20 95 ms
two-stage, k = 50 215 ms
two-stage, k = 100 415 ms
cross-encoder on all N 80,000 ms
speed-up at k = 50: 372x- Linear reranking cost: Every additional candidate contributes one further cross-encoder inference, so doubling approximately doubles the reranking latency.
- Independent of index size: The reranking term depends exclusively on , so the index can expand from thousands to millions of chunks without increasing that component.
- Recall constraint: A smaller candidate count is faster, but a relevant passage positioned outside the first candidates can never be recovered by the reranker.
In practice, is chosen as the smallest candidate count at which recall@k on the evaluation set stops improving, which keeps latency low without losing the correct passage.
Example: Reranking Orders API Passages in Python
The program below retrieves three candidates with a loose word-frequency count, then rescores them with a stricter function that rewards exact phrases and a matching section title.
import re
# Chunks from the Orders API docs: (section title, text).
chunks = {
"error-codes": ("Error codes", "Every error returns a JSON error object. Error 401 means an invalid key, error 404 means a missing order and error 429 means too many requests."),
"rate-limits": ("Rate limits", "Each API key may send 100 requests per minute. After a rate limit error, wait for the number of seconds in the Retry-After header."),
"webhooks": ("Webhooks", "A failed webhook delivery is retried after 1, 5 and 30 minutes. Your endpoint must return status 200."),
"pagination": ("Pagination", "GET /orders returns 50 orders per page. Send the cursor value to request the next page."),
}
STOP = {"a", "an", "the", "is", "to", "do", "what", "should", "after", "of", "in", "for", "and", "per", "my"}
def words(text):
return [w for w in re.findall(r"[a-z0-9-]+", text.lower()) if w not in STOP]
def first_pass(query, text):
# Stage 1: count every occurrence of a query word (fast, loose).
q = set(words(query))
return sum(1 for w in words(text) if w in q)
def rerank(query, title, text):
# Stage 2: reward exact phrases and a match on the section title.
q, t = words(query), " ".join(words(text))
phrases = sum(1 for a, b in zip(q, q[1:]) if f"{a} {b}" in t)
title_hit = all(w.rstrip("s") in q for w in words(title))
return 3 * phrases + (5 if title_hit else 0) + len(set(q) & set(t.split()))
query = "What should my client do after a rate limit error?"
stage1 = sorted(chunks, key=lambda k: (-first_pass(query, chunks[k][1]), k))[:3]
print("Stage 1 (top 3 by word count):")
for k in stage1:
print(f" {k:<12} {first_pass(query, chunks[k][1])}")
stage2 = sorted(stage1, key=lambda k: (-rerank(query, *chunks[k]), k))
print("Stage 2 (reranked):")
for k in stage2:
print(f" {k:<12} {rerank(query, *chunks[k])}")
print("Passage sent to the model:", stage2[0])Output:
Stage 1 (top 3 by word count):
error-codes 5
rate-limits 3
pagination 0
Stage 2 (reranked):
rate-limits 14
error-codes 1
pagination 0
Passage sent to the model: rate-limits- Incorrect initial ranking: The loose count places Error codes first, because the word "error" appears five times in that passage.
- Corrected ordering: The reranker detects the phrases "rate limit" and "limit error" and the matching Rate limits title, so the relevant passage moves to first position.
- Production equivalent: Real systems replace the hand-written function with a cross-encoder model, but the two-stage architecture remains identical.
Bi-Encoder Retrieval vs Cross-Encoder Reranking
| Aspect | Bi-encoder (first stage) | Cross-encoder (reranker) |
|---|---|---|
| Input | Question and passage encoded separately | Question and passage processed together |
| Speed | Fast, uses a precomputed index | Slower, one model inference per pair |
| Scale | Millions of chunks | Tens to hundreds of candidates |
| Precision | Good recall, approximate ordering | Precise ordering at the top of the list |
A reranker is most valuable when evaluation shows that the relevant passage usually appears within the top 20 candidates but is rarely ranked first. In that situation, recall is adequate and ranking precision is the limitation, which is the specific problem a second-stage relevance model addresses.
Applications of Reranking in RAG
- Developer documentation assistants: Placing the exact API section first for questions about rate limits, errors or webhooks.
- Customer support retrieval: Ranking the applicable policy passage above loosely related help articles.
- Legal and compliance search: Promoting the contract clause that contains a precise regulatory phrase.
- Source code retrieval: Ordering candidate snippets by similarity to a function name and an error message.
- Agentic retrieval: Allowing an agent in agentic RAG to retain only the strongest evidence before its next action.
Advantages
- Improved precision: The first passage the model reads is more frequently the correct one.
- Compact prompts: Fewer, more relevant passages consume less of the context window.
- Reduced hallucination: More relevant evidence lowers the probability of LLM hallucination.
- Simple integration: A reranker attaches to any existing retriever without rebuilding the index.
Limitations
- Additional latency: Every candidate requires a separate scoring operation, which slows each query.
- Recall limitation: A reranker cannot recover a passage that the first-stage retriever never returned.
- Operational cost: Hosted reranking models and LLM rerankers are billed per request.
- Parameter tuning: The candidate count and the cut-off require testing with RAG evaluation metrics.
Reranking is one component of advanced pipelines; the other architectures are compared in types of RAG.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. Why does reranking in RAG run only on the top candidates instead of on every chunk in the index?
Frequently Asked Questions
What is a RAG reranker?
A reranker is a second scoring model that reorders the passages returned by the first retriever. It reads the question and each passage together, so it judges relevance more precisely. Only the top passages after reranking go into the prompt.
Is reranking always needed in a RAG system?
Reranking is not always needed. A small, clean document set often ranks well with vector search alone. It becomes useful when the right passage is often retrieved but not ranked first.
What is the difference between a bi-encoder and a cross-encoder?
A bi-encoder turns the question and each passage into separate embeddings and compares them, which is fast and works with an index. A cross-encoder reads the question and passage as one input and outputs a score, which is slower but more accurate.
How many passages should be reranked?
Teams usually rerank somewhere between 20 and 100 candidates and keep the top 3 to 5. The right numbers depend on latency limits and are chosen by measuring retrieval quality on a test set.
Can an LLM be used as a reranker?
An LLM can score or sort passages when given the question and a list of candidates. It is often accurate but adds cost and latency, so it suits small candidate lists or offline evaluation.
Related Articles
- How RAG WorksHow RAG works in detail: the offline indexing pipeline, the online query pipeline, a traced Python example over API docs, and where each stage can fail.
- Vector Database in RAGVector database in RAG explained: how embeddings are stored, indexed and searched with cosine similarity, ANN indexes, metadata filters and a Python demo.
- Hybrid Search (BM25 + Vector)Hybrid search explained: how BM25 keyword results and vector results are merged with reciprocal rank fusion, with a runnable Python runbook search example.
- Types of RAG (Naive, Advanced, Graph RAG, Agentic RAG)Types of RAG compared: naive, advanced, modular, graph RAG and agentic RAG, how each retrieves context, typical uses, and how to choose the right design.
- RAG Evaluation MetricsRAG evaluation metrics explained: hit rate, recall@k, MRR, faithfulness and answer relevance, with a Python example on a labelled test set and its limits.
- Build a RAG Chatbot in PythonBuild a RAG chatbot in Python: chunk API docs, retrieve the closest passage, build a grounded prompt, then call Claude with the Anthropic SDK. Full code.