RAG Evaluation Metrics
RAG evaluation is the process of measuring how well a retrieval-augmented generation system finds relevant passages and how faithfully its answers follow them. It uses a labelled test set and separate metrics for the retrieval stage and the generation stage, so each failure can be traced to its source.
- Retrieval metrics: Measurements such as hit rate, recall@k (recall at k) and mean reciprocal rank that score the ranked list of retrieved chunks.
- Generation metrics: Measurements such as faithfulness and answer relevance that score the final answer.
- Labelled test set: Realistic questions paired with the chunk that answers them and, optionally, a verified reference answer.
- Faithfulness metric: The proportion of answer statements that the retrieved passages support, also called groundedness.
- LLM-as-a-judge: A language model prompted to grade answers against a rubric when exact matching is impossible.
For example, a team tests its Orders API documentation assistant on questions about authentication, rate limits and pagination, and checks both the retrieved sections and the answer sentences.
Key Characteristics of RAG Evaluation
- Component separation: Retrieval and generation are scored independently, because each stage of how RAG works fails for different reasons.
- Ground-truth labels: Retrieval metrics require the identity of the relevant chunk for every test question.
- Rank sensitivity: Metrics such as MRR reward systems that place the relevant chunk near the top of the list.
- Reference-independent scoring: Faithfulness compares the answer with the retrieved context, so a manually written reference answer is not necessary.
- Regression testing: The same test set is rerun after every change to chunking, embeddings, prompts or models.
How RAG Evaluation Works
The steps below show how to evaluate RAG as a repeatable process, from the test set to the version comparison.
- Test set construction: Collect 50 to 200 representative questions from real users, support tickets or documentation owners, and label the relevant chunk for each.
- Retrieval execution: Run every question through the retriever and record the ranked chunk identifiers for later analysis.
- Retrieval scoring: Compute hit rate@k, recall@k and MRR from the recorded rankings and the labels.
- Answer generation: Generate an answer for each question from its retrieved context.
- Answer scoring: Evaluate faithfulness and relevance, using lexical overlap rules, an LLM judge or human review.
- Comparison: Compare the scores with the previous version and keep only changes that improve them.
Common RAG Evaluation Metrics
| Metric | Stage | What it measures |
|---|---|---|
| Hit rate@k | Retrieval | Share of questions with a relevant chunk in the top k |
| Recall@k | Retrieval | Share of all relevant chunks that appear in the top k |
| Precision@k | Retrieval | Share of the top k chunks that are relevant |
| MRR | Retrieval | Mean of 1/rank of the first relevant chunk |
| Faithfulness | Generation | Share of answer statements supported by the context |
| Answer relevance | Generation | Whether the answer addresses the question asked |
Computing Retrieval Metrics
Retrieval metrics compare the ranked list that the retriever returns with the set of chunks labelled as relevant. Precision and recall ignore the order inside the top , whereas MRR and nDCG reward systems that place relevant chunks near the top:
- : the set of chunks labelled as relevant for one question.
- : the first chunks returned by the retriever.
- : the set of test questions, and : the position of the first relevant chunk for question .
- : the graded relevance of the chunk at position , for example 2 for a direct answer, 1 for partial help and 0 for an irrelevant chunk.
- : the position discount, which equals 1 at the first position and grows slowly afterwards.
- : the DCG of the ideal ordering, obtained by sorting the labelled chunks from most to least relevant.
In words, precision@k reports what proportion of the retrieved chunks is relevant, and recall@k reports what proportion of the relevant chunks was retrieved. The MRR metric averages the reciprocal of the first relevant position across all questions. The nDCG formula divides every relevance grade by a discount that increases with position, then compares the total with the best achievable total, so a perfect ordering scores exactly 1.
The worked example evaluates one Orders API question, "How should a client retry after a 429 response?", where the rate-limits chunk answers it directly (grade 2) and the error-codes chunk helps partially (grade 1). The retriever returns error-codes, pagination, rate-limits and webhooks:
- Precision and recall: Two of the first three retrieved chunks are relevant, and both labelled chunks appear within the top three, so precision equals 0.67 and recall is perfect.
- Reciprocal rank: The partially relevant error-codes chunk occupies the first position, so the reciprocal rank for this individual question equals 1.
- Discounted gain: The direct answer is positioned third, where its relevance grade of 2 is divided by a logarithmic discount of 2, so it contributes only 1.
- Ideal ordering: The optimal arrangement places rate-limits first and error-codes second, and the ratio between the actual and ideal totals produces the normalised score.
import math
# "How should a client retry after a 429 response?"
# Graded labels: 2 = answers the question, 1 = partly useful, 0 = not relevant.
relevance = {"rate-limits": 2, "error-codes": 1}
ranked = ["error-codes", "pagination", "rate-limits", "webhooks"]
k = 3
top = ranked[:k]
hits = sum(doc in relevance for doc in top)
print(f"precision@{k} = {hits}/{k} = {hits / k:.3f}")
print(f"recall@{k} = {hits}/{len(relevance)} = {hits / len(relevance):.3f}")
first = next(i for i, doc in enumerate(ranked, start=1) if doc in relevance)
print(f"reciprocal rank = 1/{first} = {1 / first:.3f}")
def dcg(gains):
return sum(rel / math.log2(i + 1) for i, rel in enumerate(gains, start=1))
gains = [relevance.get(doc, 0) for doc in top]
ideal = sorted(relevance.values(), reverse=True)[:k]
for i, rel in enumerate(gains, start=1):
print(f" rank {i}: rel {rel} / log2({i + 1}) = {rel / math.log2(i + 1):.3f}")
print(f"DCG@{k} = {dcg(gains):.3f}, IDCG@{k} = {dcg(ideal):.3f}")
print(f"nDCG@{k} = {dcg(gains) / dcg(ideal):.3f}")precision@3 = 2/3 = 0.667
recall@3 = 2/2 = 1.000
reciprocal rank = 1/1 = 1.000
rank 1: rel 1 / log2(2) = 1.000
rank 2: rel 0 / log2(3) = 0.000
rank 3: rel 2 / log2(4) = 1.000
DCG@3 = 2.000, IDCG@3 = 2.631
nDCG@3 = 0.760- Hidden ordering problem: Recall and reciprocal rank both appear perfect, yet an nDCG@3 of 0.76 reveals that the most informative chunk is buried at the third position.
- Graded relevance: Only nDCG distinguishes a chunk that genuinely answers the question from a chunk that merely mentions status 429.
- Averaging: MRR and nDCG are averaged across every question in the evaluation set, as the Python example in the next section demonstrates for MRR over four questions.
In practice, nDCG is the metric that detects whether reranking in RAG genuinely moves the best chunk to the top, which precision and recall cannot measure.
Example: Retrieval Metrics and Faithfulness in Python
The program below scores logged retrieval results for four labelled questions, then checks each sentence of a generated answer against its source passage.
import re
# Labelled set: each question, the chunk that answers it, and the
# retriever's ranked results for that question.
tests = [
("How do I authenticate to the Orders API?", "authentication", ["authentication", "error-codes", "rate-limits"]),
("How many requests per minute are allowed?", "rate-limits", ["pagination", "rate-limits", "webhooks"]),
("How do I get the next page of orders?", "pagination", ["pagination", "webhooks", "authentication"]),
("What does error 404 mean?", "error-codes", ["webhooks", "authentication", "pagination"]),
]
def hit_rate(k):
# Share of questions whose relevant chunk is in the top k results.
return sum(rel in ranked[:k] for _, rel, ranked in tests) / len(tests)
def mrr():
# Mean of 1/rank of the relevant chunk (0 when it is missing).
return sum(1 / (ranked.index(rel) + 1) if rel in ranked else 0 for _, rel, ranked in tests) / len(tests)
print(f"Hit rate@1: {hit_rate(1):.2f}")
print(f"Hit rate@3: {hit_rate(3):.2f}")
print(f"MRR: {mrr():.3f}")
# Faithfulness: is each answer sentence supported by the passage?
passage = "Each API key may send 100 requests per minute. Extra requests return status 429 with a Retry-After header."
answer = "You can send 100 requests per minute per API key. Extra requests return status 429. Limits reset every hour."
STOP = {"a", "an", "the", "you", "can", "per", "with", "every", "each", "may"}
def terms(text):
return {w for w in re.findall(r"[a-z0-9-]+", text.lower()) if w not in STOP}
source = terms(passage)
supported = 0
sentences = re.split(r"(?<=\.)\s+", answer)
for sentence in sentences:
share = len(terms(sentence) & source) / len(terms(sentence))
ok = share >= 0.75
supported += ok
print(f"{'SUPPORTED' if ok else 'NOT FOUND'} ({share:.2f}): {sentence}")
print(f"Faithfulness: {supported}/{len(sentences)} sentences supported")Output:
Hit rate@1: 0.50
Hit rate@3: 0.75
MRR: 0.625
SUPPORTED (1.00): You can send 100 requests per minute per API key.
SUPPORTED (1.00): Extra requests return status 429.
NOT FOUND (0.00): Limits reset every hour.
Faithfulness: 2/3 sentences supported- Retrieval diagnosis: Hit rate rises from 0.50 at k=1 to 0.75 at k=3, so the rate-limits chunk is found but ranked second; the error-codes question fails completely. With one relevant chunk per question, recall@k equals hit rate@k.
- Ranking quality: MRR is 0.625, the average of the individual reciprocal ranks (1 + 0.5 + 1 + 0) / 4.
- Unsupported statement: "Limits reset every hour" has no support in the passage, so the faithfulness check flags it as a probable LLM hallucination.
Applications of RAG Evaluation
- Documentation assistants: Verifying that API answers quote the correct section before a release.
- Chunking experiments: Comparing chunking strategies by their effect on recall@k.
- Reranker selection: Measuring whether reranking in RAG raises MRR enough to justify its latency.
- Model upgrades: Confirming that a replacement language model remains faithful to the retrieved context.
- Production monitoring: Sampling real conversations from production and evaluating them every week.
- Agent testing: Extending the same approach to AI agent evaluation, where several retrieval steps occur per task.
Advantages
- Precise diagnosis: Separate metrics show whether retrieval or generation caused a wrong answer.
- Objective comparison: Numeric scores replace impressions when choosing between designs from types of RAG.
- Early detection: Regressions appear in automated test results before developers or customers report them.
- Continuous integration: Retrieval metrics are inexpensive to compute automatically on every code change.
Limitations
- Labelling effort: Identifying the relevant chunk for each question requires domain knowledge and time.
- Judge reliability: LLM judges can be inconsistent and require periodic comparison with human ratings.
- Lexical limitations: Word-overlap faithfulness misses paraphrases and accepts statements that copy vocabulary but change the meaning.
- Test set drift: Labels become outdated when documents are rechunked or rewritten.
- Incomplete coverage: A high score on a small test set does not guarantee reliable quality on unusual questions.
Evaluation libraries such as Ragas and DeepEval implement these metrics with LLM judges. The complete pipeline being evaluated is built step by step in build a RAG chatbot in Python.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In RAG evaluation, a relevant chunk ranked third contributes what value to MRR for that question?
Frequently Asked Questions
What are the most important metrics for evaluating RAG?
Most teams track a retrieval metric, such as recall@k or MRR, and a generation metric, such as faithfulness. Answer relevance is often added to check that the answer addresses the question.
What is faithfulness in RAG evaluation?
Faithfulness is the share of statements in an answer that the retrieved passages support. A low score means the model added facts that are not in the context, which is a form of hallucination.
How many questions does a RAG test set need?
A set of 50 to 200 realistic questions is a common starting point. It should cover frequent questions, rare edge cases and questions the documents cannot answer.
What is the difference between recall@k and hit rate@k?
Hit rate@k counts a question as a success if any relevant chunk is in the top k. Recall@k measures the fraction of all relevant chunks found, so the two are equal when each question has exactly one relevant chunk.
Can an LLM evaluate RAG answers?
An LLM can act as a judge by grading answers against the context and a rubric. Its scores should be checked against human ratings on a sample, because judges can be inconsistent.
Related Articles
- How RAG WorksHow RAG works in detail: the offline indexing pipeline, the online query pipeline, a traced Python example over API docs, and where each stage can fail.
- Reranking in RAGReranking in RAG explained: how a second, stricter scorer reorders retrieved passages, cross-encoders vs bi-encoders, a Python example and the trade-offs.
- Chunking Strategies in RAGChunking strategies in RAG explained: fixed-size, heading-based and semantic chunking, how to choose chunk size and overlap, with a Python comparison.
- LLM Hallucination: Causes and FixesLearn what LLM hallucination is, why models invent facts and code APIs, and how to reduce it with RAG and validation, with a runnable Python checker.
- AI Agent EvaluationAI agent evaluation explained: task success, tool-call accuracy, trajectory checks, cost, latency and regression suites, with a Python trajectory scorer.
- Build a RAG Chatbot in PythonBuild a RAG chatbot in Python: chunk API docs, retrieve the closest passage, build a grounded prompt, then call Claude with the Anthropic SDK. Full code.