What Is RAG Evaluation? A Developer Guide to Measuring Retrieval-Augmented Generation Quality

TLDR: RAG evaluation measures two things separately: whether retrieval puts the right evidence into the model's context, and whether the answer stays faithful to that evidence and addresses the question. Score retrieval with recall@k, MRR, and nDCG against labeled chunks. Score generation with faithfulness, answer relevance, and context precision and recall, usually through an LLM judge. Then version the test set and gate every change in CI.
This guide covers measurement. For construction, see how to build RAG with web search; for the pattern itself, what retrieval-augmented generation is; for choosing a retrieval layer, API for RAG. The goal is a scorecard that shows which stage broke when quality drops, plus code that runs offline before any judge model or paid API is involved.
What does RAG evaluation measure?
A RAG answer can fail in three places. Retrieval can miss the evidence. Ranking can push it below the cutoff you pass to the model. Generation can receive good evidence and still add unsupported claims or answer a different question. End-to-end correctness tells you something failed; component metrics tell you where.
| Failure | Symptom | Metric that catches it | Label you need |
|---|---|---|---|
| Relevant chunk never retrieved | Confident answer from model memory, or a refusal | Recall@k, success@k | Relevant chunk IDs per query |
| Relevant chunk ranked below the cutoff | Answer misses a fact your index contains | MRR, nDCG@k, context precision | Graded relevance, or a reference answer |
| Noise in the context | Answer mixes in off-topic facts | Context relevance, noise sensitivity | Judge model |
| Unsupported claims | Hallucination despite good context | Faithfulness or groundedness | Judge model plus the retrieved context |
| Grounded but off-question | Accurate text that does not answer | Answer relevance | Judge model |
| Wrong final answer | Disagrees with the known answer | Answer or factual correctness | Reference answer |
Retrieval metrics are arithmetic over labels, cheap enough for every commit. Generation metrics mostly need an LLM judge, so they are slower, costlier, and noisier. That split shapes what you label and what blocks a merge.
Retrieval metrics: recall@k, MRR, and nDCG
Retrieval metrics compare a ranked list of chunk IDs with relevance judgments, often called qrels: the IDs a person marked relevant for each query, ideally graded, such as 2 for "answers the question" and 1 for "partially relevant." Set k to the number of chunks you pass to the model, since nothing below that cutoff reaches generation.
Recall@k is the fraction of a query's relevant documents that appear in the top k. The ir_measures reference flags a naming trap: some tasks use "recall@k" to mean whether any relevant document appears in the top k, a measure TREC convention calls Success@k. Check which one your tool reports before you compare numbers across tools. Make recall@k your primary retrieval gate: a chunk that is never retrieved cannot be cited.
MRR, mean reciprocal rank, averages 1 divided by the rank of the first relevant result, or 0 when none comes back. It suits lookups where one passage holds the answer; MS MARCO's passage ranking task, built around ranking the most relevant passage as high as possible, is evaluated with MRR. It ignores everything after the first hit, so it undercounts questions that need several chunks.
nDCG@k uses graded labels and discounts gain by position: the Stanford IR book scores each result as (2**grade - 1) / log2(1 + rank), sums the top k, and normalizes so a perfect ranking scores 1. Position matters because models do not read context evenly: the Lost in the Middle study found performance is often highest when relevant information sits at the start or end of the context, and degrades when it sits in the middle.
Two details change the numbers. Compute the ideal ranking from every judged document, not only what the retriever returned, or a retriever that misses the strongest chunk can still score 1.0. And skip precision@k as a gate: the IR book calls it the least stable of the common measures, and it averages poorly because relevant-document counts vary by query.
Score retrieval offline with the Python standard library
The script needs Python 3.9 or later and no packages, credentials, or network access. Save the three JSON fixtures and two Python files in one directory. The fixtures are synthetic: four help-center queries, where q2 is a retrieval miss, q4 has no relevant document, and a grade of 0 records a document someone judged and rejected.
{
"q1": {"refund-policy#2": 2, "refund-policy#3": 1},
"q2": {"sla-uptime#1": 2},
"q3": {"sso-setup#4": 2, "sso-setup#5": 2, "okta-guide#1": 1},
"q4": {"pricing#3": 0}
}
Save that as qrels.json. The run file, run.json, holds the top five chunk IDs your retriever returned for each query, in rank order:
{
"q1": ["shipping#1", "refund-policy#2", "refund-policy#9", "refund-policy#3", "returns-faq#1"],
"q2": ["pricing#3", "status-page#1", "sla-uptime#7", "support-hours#2", "sla-uptime#2"],
"q3": ["sso-setup#5", "okta-guide#1", "sso-setup#1", "scim#2", "sso-setup#4"],
"q4": ["pricing#3", "pricing#1", "plans#2", "billing#4", "invoices#1"]
}
The baseline, baseline.json, holds the mean metrics from the last run you accepted, committed next to the test set:
{"recall@5": 0.70, "success@5": 0.6667, "mrr@5": 0.55, "ndcg@5": 0.52}
Save the scorer as rag_eval.py:
"""Retrieval metrics for RAG evaluation: recall@k, success@k, MRR, and nDCG@k.
Python 3.9+ standard library only; no network calls.
Usage: python3 rag_eval.py qrels.json run.json --k 5 [--baseline baseline.json]
"""
import argparse
import json
import math
import sys
def top_k(ranking, k):
"""First k unique IDs in rank order, so a duplicated chunk cannot count twice."""
seen, out = set(), []
for doc_id in ranking:
if doc_id not in seen:
seen.add(doc_id)
out.append(doc_id)
if len(out) == k:
break
return out
def recall_at_k(ranking, grades, k):
relevant = {d for d, g in grades.items() if g > 0}
return sum(1 for d in top_k(ranking, k) if d in relevant) / len(relevant)
def success_at_k(ranking, grades, k):
return float(any(grades.get(d, 0) > 0 for d in top_k(ranking, k)))
def reciprocal_rank(ranking, grades, k):
for rank, d in enumerate(top_k(ranking, k), start=1):
if grades.get(d, 0) > 0:
return 1.0 / rank
return 0.0
def dcg(gains):
return sum((2 ** g - 1) / math.log2(rank + 1) for rank, g in enumerate(gains, start=1))
def ndcg_at_k(ranking, grades, k):
got = [grades.get(d, 0) for d in top_k(ranking, k)]
# The ideal ranking comes from every judged document, not only the retrieved ones.
ideal = sorted((g for g in grades.values() if g > 0), reverse=True)[:k]
return dcg(got) / dcg(ideal)
METRICS = {"recall": recall_at_k, "success": success_at_k,
"mrr": reciprocal_rank, "ndcg": ndcg_at_k}
def evaluate(qrels, run, k):
if not isinstance(k, int) or k < 1:
raise ValueError("k must be a positive integer")
unexpected = sorted(set(run) - set(qrels))
if unexpected:
raise ValueError("run has queries that are not in qrels: %s" % unexpected)
per_query, unscored = {}, []
for qid in sorted(qrels):
grades = qrels[qid]
if not all(isinstance(g, int) and g >= 0 for g in grades.values()):
raise ValueError("%s: grades must be integers of 0 or more" % qid)
if not any(g > 0 for g in grades.values()):
unscored.append(qid) # nothing relevant exists: score abstention instead
continue
ranking = run.get(qid, []) # a query the retriever skipped scores zero
if not isinstance(ranking, list):
raise ValueError("%s: ranking must be a list of IDs" % qid)
per_query[qid] = {name: fn(ranking, grades, k) for name, fn in METRICS.items()}
if not per_query:
raise ValueError("no query has a relevant document")
mean = {"%s@%d" % (name, k): sum(s[name] for s in per_query.values()) / len(per_query)
for name in METRICS}
return {"k": k, "queries": len(per_query), "unscored": unscored,
"mean": mean, "per_query": per_query}
def regressions(current, baseline, max_drop):
"""Baseline metrics that are missing or fell more than max_drop."""
bad = {}
for name, was in baseline.items():
now = current.get(name)
if now is None or was - now > max_drop:
bad[name] = (was, now)
return bad
def main(argv=None):
parser = argparse.ArgumentParser(description="Score a retrieval run against qrels.")
parser.add_argument("qrels")
parser.add_argument("run")
parser.add_argument("--k", type=int, default=5)
parser.add_argument("--baseline", help="JSON of mean metrics from the last accepted run")
parser.add_argument("--max-drop", type=float, default=0.02)
args = parser.parse_args(argv)
with open(args.qrels) as fh:
qrels = json.load(fh)
with open(args.run) as fh:
run = json.load(fh)
report = evaluate(qrels, run, args.k)
print(json.dumps({"mean": report["mean"], "unscored": report["unscored"]},
indent=2, sort_keys=True))
if not args.baseline:
return 0
with open(args.baseline) as fh:
bad = regressions(report["mean"], json.load(fh), args.max_drop)
for name, (was, now) in sorted(bad.items()):
print("REGRESSION %s: baseline %s, now %s" % (name, was, now), file=sys.stderr)
return 1 if bad else 0
if __name__ == "__main__":
sys.exit(main())
Save the tests as test_rag_eval.py. Each expected value was worked out by hand from the fixtures:
"""Offline checks for rag_eval.py. Run: python3 -m unittest test_rag_eval.py"""
import json
import math
import unittest
from rag_eval import evaluate, ndcg_at_k, regressions
with open("qrels.json") as fh:
QRELS = json.load(fh)
with open("run.json") as fh:
RUN = json.load(fh)
class RetrievalMetricTests(unittest.TestCase):
def test_values_match_hand_calculation(self):
q = evaluate(QRELS, RUN, 5)["per_query"]
self.assertAlmostEqual(q["q1"]["recall"], 1.0)
self.assertAlmostEqual(q["q1"]["mrr"], 0.5)
dcg = 3 / math.log2(3) + 1 / math.log2(5)
ideal = 3 / math.log2(2) + 1 / math.log2(3)
self.assertAlmostEqual(q["q1"]["ndcg"], dcg / ideal)
self.assertEqual(q["q2"], {"recall": 0.0, "success": 0.0, "mrr": 0.0, "ndcg": 0.0})
def test_ideal_ranking_uses_every_judged_document(self):
# At k=3 the run misses sso-setup#4 (grade 2). Normalizing by the
# retrieved list alone would report a perfect 1.0 here.
self.assertLess(ndcg_at_k(RUN["q3"], QRELS["q3"], 3), 0.7)
def test_duplicate_chunks_do_not_inflate_scores(self):
ranking = ["sla-uptime#1", "sla-uptime#1", "sla-uptime#1"]
self.assertAlmostEqual(ndcg_at_k(ranking, QRELS["q2"], 3), 1.0)
def test_skipped_query_scores_zero_and_stays_in_the_mean(self):
partial = {q: r for q, r in RUN.items() if q != "q2"}
report = evaluate(QRELS, partial, 5)
self.assertEqual(report["queries"], 3)
self.assertEqual(report["mean"], evaluate(QRELS, RUN, 5)["mean"])
def test_query_without_relevant_documents_is_reported_not_scored(self):
self.assertEqual(evaluate(QRELS, RUN, 5)["unscored"], ["q4"])
def test_gate_flags_drops_beyond_tolerance(self):
with open("baseline.json") as fh:
bad = regressions(evaluate(QRELS, RUN, 5)["mean"], json.load(fh), 0.02)
self.assertEqual(sorted(bad), ["mrr@5", "recall@5"])
if __name__ == "__main__":
unittest.main()
Run the tests, then the gate:
python3 -m unittest test_rag_eval.py
python3 rag_eval.py qrels.json run.json --k 5 --baseline baseline.json --max-drop 0.02
echo "gate exit code: $?"
All six tests pass. The scorer reports recall@5 of 0.667, MRR@5 of 0.5, and nDCG@5 of 0.509, and exits 1 because recall and MRR fell more than 0.02 below the baseline. The tests pin four behaviors a hand-rolled scorer can get wrong: q3 at k=3 scores 0.673, not the 1.0 that retrieved-list normalization reports; duplicates cannot push nDCG above 1; a skipped query scores zero instead of vanishing from the average; and q4 is reported as unscored, to be graded on abstention.
Generation metrics in Ragas, DeepEval, and TruLens
Generation metrics ask whether the answer is supported by the context, whether it addresses the question, and whether the context was precise and complete. Ragas, DeepEval, and TruLens share metric names but not always definitions. The table follows each project's docs for its current PyPI release as of September 2026: Ragas 0.4.3, DeepEval 4.2.6, and TruLens 2.14.0.
| Question | Ragas | DeepEval | TruLens |
|---|---|---|---|
| Is the answer supported by the context? | Faithfulness: share of response claims that can be inferred from the retrieved context | FaithfulnessMetric: share of claims that do not contradict the retrieval context | Groundedness: each claim in the response checked for supporting evidence in the context |
| Does the answer address the question? | Answer relevancy: mean embedding similarity between the user input and questions generated from the response (3 by default) | AnswerRelevancyMetric: share of output statements relevant to the input | Answer relevance: relevance of the final response to the user input |
| Is the context precise? | Context precision: rank-weighted precision of chunks judged against a reference (or the response, as context utilization) | ContextualPrecisionMetric: weighted cumulative precision against the expected output; ContextualRelevancyMetric: share of relevant context statements | Context relevance: each chunk scored for relevance to the query, then aggregated |
| Is the context complete? | Context recall: share of reference-answer claims supported by the retrieved context | ContextualRecallMetric: share of expected-output statements attributable to the retrieval context | Not part of the RAG triad, which needs no reference answer |
Read the faithfulness row twice. Under the Ragas definition, a plausible claim the context does not support lowers the score; under DeepEval's, it counts as truthful unless it contradicts the context. Confirm the definition before comparing anyone's faithfulness number with yours. Ragas also notes that its embedding-based answer relevancy is not guaranteed to stay between 0 and 1.
Reference answers decide what you can measure. Ragas's LLM-based context recall and DeepEval's contextual precision and recall need one; the TruLens RAG triad of context relevance, groundedness, and answer relevance does not, which makes it a practical start before labeling. To pick a framework by layer (completions, RAG, or agents), see the LLM evaluation framework guide.
Pin the library version. Judge prompts ship inside the package (DeepEval exposes them as overridable evaluation templates), so an upgrade is a judge change that needs a fresh baseline. Ragas's docs also slate its legacy metrics API for removal in version 1.0.
LLM-as-judge caveats that move your scores
An LLM judge is a model with its own error profile. Zheng et al. (2023) documented position, verbosity, and self-enhancement biases and limited reasoning ability, and also found that strong judges such as GPT-4 reached over 80% agreement with human preferences, the level humans reach with each other. Judges are usable at scale, and they can be wrong in consistent directions. When a judge compares two answers, run each comparison in both orders; the deep research evaluation guide shows that control in a working harness.
Alignment belongs to a judge configuration, not a model. The TruLens judge alignment guide notes that agreement varies by model, dataset, evaluated property, and annotator expertise, recommends a human-labeled golden set split into development, validation, and held-out portions, and states a rule worth adopting: "Do not use an LLM to create the labels used to prove that the same judge aligns."
- Pin the judge. Fix model, version, temperature, and prompt, and store them with every score. DeepEval's docs list gpt-5.4 as the default judge; passing
modelexplicitly keeps an upgrade from changing it. - Use a different model family from your generator where you can, since self-enhancement bias is a judge favoring answers it generated.
- Re-grade a frozen sample every run. If scores on unchanged answers move, the judge drifted, and that run's regressions are suspect.
- Expect flips at the threshold. DeepEval's CI docs note that LLM evals are non-deterministic and borderline cases flip between runs; marking them flaky makes them warn instead of fail.
- Correct the judge with a few human labels. ARES fine-tunes lightweight judges and uses a small human-annotated set for prediction-powered inference, reporting accurate evaluation with only a few hundred human annotations.
Build and version the test set
Draw questions from production logs (with personal data removed), support tickets, and subject-matter experts, and tag each with a slice. The Ragas test set guide separates single-hop from multi-hop and specific from abstract queries. Add time-sensitive, conflicting-source, and unanswerable slices, and grade the unanswerable ones on abstention rather than recall.
Each item needs a stable ID, the question, graded relevance labels, a reference answer, and an answer-validity date if it is time-sensitive. Attach labels to a unit that survives re-chunking, such as a document ID plus a passage offset; labels on raw chunk IDs are orphaned the moment chunk size changes. Ragas builds synthetic test sets from your documents with a knowledge-graph approach, and DeepEval's Synthesizer generates goldens from documents. Review generated items, and report synthetic and real slices separately, because a question written from a chunk can reuse its wording and flatter your retriever.
Version the set like code. Store it as JSONL in the repository, never edit a released version in place, record a checksum, and write the dataset version, index build, embedding model, generator, and judge configuration into every results file. Keep a held-out slice you never tune against.
Size follows the difference you need to detect. For a pass or fail metric near 80%, the 95% margin of error is about plus or minus 8 points on 100 questions and 4 points on 400 (normal approximation to the binomial). Paired comparisons are tighter, since only questions that flip carry signal, but a 3-point drop on 100 questions is still just three answers. Randomness in AI Benchmarks covers the statistics in more depth.
Watch for unjudged results. When a new embedding model surfaces chunks nobody labeled, the scorer counts them as irrelevant, so a real improvement can look like a regression. Label the new top-k before deciding; ir_measures offers a judged_only option on nDCG for this case.
Gate regressions in CI
Split the suite by cost and determinism, and give each tier its own trigger.
- Every pull request: retrieval metrics. Run
rag_eval.pyagainst a frozen index and fail the job on a nonzero exit. With a frozen index and a deterministic retriever the numbers repeat exactly, so any drop comes from a code, chunking, or embedding change, and the tolerance (0.02 here) only sets how large a real drop can merge without review. Updatebaseline.jsononly in the pull request that earns the new number. - Pull requests that touch prompts, models, or context assembly: judged generation metrics. Replay stored retrieval results so only generation varies, and gate with thresholds.
- On a schedule: live end-to-end runs with fresh retrieval, repeated trials, and a comparison with the last accepted run.
For the second tier, DeepEval plugs into pytest. Per its CI documentation, when a metric in assert_test() falls below its threshold under deepeval test run, the build goes red, while evaluate() collects results without failing anything. The example replays frozen chunks. Its thresholds are placeholders to set from your baseline, and it was syntax-checked but not run, since it needs the library and a judge API key.
# test_rag_generation.py for deepeval==4.2.6. Run: deepeval test run test_rag_generation.py
import json
import pytest
from deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric, ContextualRecallMetric, FaithfulnessMetric
from deepeval.test_case import LLMTestCase
from my_rag import generate_answer # your generator, called on stored chunks (replay)
JUDGE = "gpt-5.4" # pin the judge explicitly and record it next to the scores
with open("rag_testset_v3.jsonl") as fh:
CASES = [json.loads(line) for line in fh if line.strip()]
@pytest.mark.parametrize("case", CASES, ids=[c["query_id"] for c in CASES])
def test_rag_answer(case):
chunks = case["retrieved_chunks"] # frozen retrieval, so only generation varies
test_case = LLMTestCase(
input=case["question"],
actual_output=generate_answer(case["question"], chunks),
expected_output=case["reference_answer"],
retrieval_context=chunks,
)
assert_test(test_case=test_case, metrics=[
FaithfulnessMetric(threshold=0.8, model=JUDGE),
AnswerRelevancyMetric(threshold=0.7, model=JUDGE),
ContextualRecallMetric(threshold=0.7, model=JUDGE),
])
Because assert_test() fails per test case, one borderline answer can block a merge. That suits high-risk slices such as pricing or compliance; for the long tail, mark cases flaky or gate on the aggregate.
The scheduled tier needs care with a web retriever, because a live index changes between runs. Store each run's URLs, ranks, and snippets, and label newly surfaced URLs before trusting a drop. The You.com Web Search API returns results under results.web and results.news at $5.00 per 1,000 calls as of September 2026, so a 500-question live suite spends $2.50 on search per run, plus $1.00 per 1,000 pages if you extract full pages crawled live. The default self-serve rate limit of 10 requests per second puts a floor of about 50 seconds on those calls. For freshness slices, note that news page_age is the article's UTC publication timestamp while web page_age is only the age of the result; neither is a crawl date.
You.com's evaluation guide makes the same end-to-end point: test search, synthesis, and grading together with your actual model and prompts, starting from default settings. To compare search providers instead, use the web search API evaluation protocol; for agents that retrieve in a loop, see AI agent evaluation.
LI Test
LI Test
Share Article:
Related resources.

What Is a Legal Research API? Building Cited Legal Research Into Applications
September 16, 2026
Blog

What Is a Price Monitoring API? How to Build One With the You.com Contents API
September 2, 2026
Blog
.png)
What Is the You.com Contents API? Clean Page Content From Any URL
September 2, 2026
Blog


