Visual guide
A Visual Guide to RAG Evaluation
A RAG answer can fail because retrieval missed the evidence, reranking buried it, context construction damaged it, or generation ignored it. A practical, diagram-led guide to measuring every layer and fixing the right one.
15 min read
A support assistant gives a confident answer about customer-data exports. The answer is wrong. The team rewrites the prompt, raises the temperature, lowers it again, swaps the model, and adds a stern instruction to “use only the supplied context.” Nothing fixes the failure consistently.
The relevant policy was in the index the whole time. It appeared at rank 23. The reranker only received the first 20 candidates.
This is the central problem with evaluating retrieval-augmented generation: the user sees one answer, but the answer is the output of several systems with different contracts. A single end-to-end score tells you whether the experience failed. It rarely tells you what to repair.
A useful RAG evaluation therefore works like an instrument panel. It measures whether the evidence entered the candidate set, whether the best evidence survived the context cutoff, whether the assembled context preserved the decisive details, whether the generated claims stayed inside that evidence, and whether the system knew when not to answer.
This guide builds that panel from first principles.
What’s in this guide: four contracts · evaluation records · retrieval · reranking · grounding · abstention · slices · diagnosis · a minimal harness · the operating loop · references
One answer, four contracts
RAG is often drawn as a two-box system: retriever, then generator. That is a helpful introduction and a poor debugging model. Production systems usually contain at least four contracts.
- Retrieval: produce a candidate set that contains the evidence needed to answer.
- Ranking: put the most useful evidence inside a much smaller context budget.
- Context construction: preserve the relevant spans, metadata, ordering, and authority signals when chunks are assembled.
- Generation: produce claims supported by that context, or abstain when support is missing.
The boxes fail differently. A retriever can miss the policy because the query says “external staff” while the document says “contractor.” A reranker can find the policy and still bury it beneath semantically similar FAQs. Context construction can cut a table away from its header or keep the permission but drop its expiry condition. The generator can receive perfect evidence and choose an answer from its parametric memory instead.
The repair follows the failed contract:
- Low retrieval recall: improve chunking, query rewriting, metadata filters, sparse-dense fusion, or the embedding model.
- Good recall but poor final ranking: improve the reranker, candidate count, feature set, or cutoff.
- Good final context but incomplete evidence: fix chunk boundaries, document parsing, deduplication, and context packing.
- Good evidence but unsupported claims: change the generation policy, prompting, claim verification, or model.
This sounds obvious written as a list. In practice, many teams record only the question and final answer. That removes the evidence needed to distinguish the four cases.
Start with an evaluation record, not a score
An evaluation set is not a spreadsheet of questions and ideal prose answers. A polished reference answer is often the least reusable part of the record: two very different responses can both be correct, and string similarity punishes the difference.
The durable unit is a small specification of what the system must know and where that knowledge lives.
For each case, store:
- Question: the language a real user would use, including ambiguity and shorthand.
- Answerability: whether the current authorized corpus contains enough evidence to answer.
- Expected claims: the smallest factual units a correct answer must contain.
- Supporting evidence: stable document and span identifiers, preferably with document version.
- Slice labels: the reason the case is difficult: paraphrase, multi-hop, temporal, permission-sensitive, unanswerable, and so on.
- Provenance: where the case came from and who last reviewed it.
Evidence identifiers matter more than a copied reference paragraph. A copied paragraph goes stale when a policy changes. A stable source identifier and version tell you that the evaluation case itself needs review.
Start with real work. Sample questions from support logs, search misses, analyst workflows, and subject-matter experts. Add synthetic cases to widen coverage, but do not let generated questions become the entire benchmark. Synthetic data tends to inherit the vocabulary and clean structure of its source documents; real users do neither.
My starting point for a new system is 30–50 carefully reviewed cases that represent the most costly failures. That is too small to declare statistical victory and large enough to expose architecture mistakes. Grow the set from production traces after that.
Retrieval: did the evidence enter the candidate set?
The retriever’s job is not to answer the question. Its job is to make the answer possible downstream.
For a case with a known set of relevant evidence items, the basic metric is recall at a cutoff:
recall@k = relevant items found in top k / all relevant items
Suppose a question needs two policy sections and the retriever finds both in its top 50. Recall@50 is 1.0. If one section is absent, recall@50 is 0.5. For cases where any one of several passages is sufficient, record that rule explicitly and use hit rate: did at least one sufficient passage appear?
The cutoff must match the actual handoff. If the reranker receives 50 candidates, recall@10 measures an imaginary system. Measure recall at the candidate boundary that exists in production.
Precision can be useful here, especially when retrieval is expensive or irrelevant documents create latency. But for the broad candidate stage, recall usually has priority. The reranker cannot promote evidence it never receives.
Three details make retrieval evaluation more honest:
Evaluate at the evidence unit the system uses. If your index stores chunks, label chunk or span identifiers rather than only document identifiers. Retrieving the right 80-page PDF is not success when the decisive paragraph never reaches the model.
Keep the authorization filter in the test. A passage retrieved without required access control is not a relevant hit. It is a security failure. Evaluation should execute the same tenant, role, and metadata filters used in production.
Record corpus version. A retrieval miss against yesterday’s index and a miss against today’s source of truth are different events. Without index and document versions, the result cannot be reproduced.
Reranking: did the right evidence reach the context?
A retriever is allowed to be generous. The final context is not. Every irrelevant chunk consumes tokens, attention, and an opportunity to distract the model.
The ranking stage therefore asks a different question: how early does useful evidence appear?
- Precision@k measures how much of the final set is relevant.
- Mean Reciprocal Rank (MRR) rewards placing the first relevant result early. It works well when one passage can answer the question.
- nDCG@k handles several results with graded relevance and discounts useful evidence that appears late.
- Context recall or coverage measures how many expected claims are supported somewhere in the final assembled context.
Do not collapse these into a universal “retrieval score.” A system can have perfect candidate recall and terrible context precision. It can also have high precision while missing one mandatory condition. For a legal or access-control answer, that missing condition can matter more than every correct sentence around it.
The cleanest experiment holds generation constant. Save the candidate set, rerank it with version A and version B, and compare the evidence metrics before running either through an LLM. This makes ranking changes cheap to test and removes generation variance from the decision.
Context construction deserves its own trace even when it is not a learned model. Store the exact text sent to the generator, in order, after deduplication and truncation. The failure may be in the glue: a markdown parser dropped a table header, two overlapping chunks repeated one policy five times, or a high-authority document lost its date.
Generation: did the answer stay inside the evidence?
Once the correct evidence reaches the prompt, the evaluation changes from search to claims.
Split the answer into atomic claims. For each claim, ask whether the supplied context supports it. Then ask whether the set of claims covers the expected answer. These are separate tests:
- Faithfulness or supportedness: What fraction of generated claims is supported by the provided context?
- Completeness: What fraction of expected claims appears in the answer?
- Answer correctness: Does the response solve the user’s task, including conditions and scope?
- Citation validity: Do the cited sources exist and correspond to retrieved evidence?
- Citation entailment: Does each cited span actually support the claim attached to it?
Citation validity is the easiest check and the weakest guarantee. A model can cite a real policy while changing its meaning. Faithfulness evaluates the relationship between claims and evidence. Correctness evaluates the relationship between the answer and the user’s need.
LLM judges are useful here because claim support is a semantic task. They are not ground truth. Calibrate them against a small human-labeled set, keep the judging rubric and model version fixed, and sample disagreements for review. The ARES paper demonstrates this hybrid pattern: automated judges become much more useful when anchored by a few hundred human annotations and statistically corrected rather than trusted in isolation.
Also run deterministic checks wherever possible. Citation identifiers, required sections, numeric values, dates, JSON schemas, and banned unsupported sources do not need a language model to score them.
Abstention is part of correctness
A RAG system has two legitimate actions: answer from evidence or decline because sufficient evidence is unavailable. Evaluation must include both answerable and unanswerable questions or it will reward the wrong behavior.
For unanswerable cases, test several causes:
- The corpus genuinely lacks the information.
- The information exists but the user is not authorized to access it.
- Two authoritative sources conflict.
- The question requires a newer document version than the index contains.
- The question is ambiguous enough that a clarifying question is safer than an answer.
Track unsupported-answer rate on unanswerable cases and over-refusal rate on answerable cases. A system can drive hallucinations to zero by refusing everything. That produces a perfect safety statistic and a useless product.
The desired behavior may be more specific than “I don’t know.” In some products the correct action is to ask a clarifying question. In others it is to open a support ticket, surface the conflicting sources, or say which document is missing. Encode the acceptable action in the case.
An average score will lie to you
An aggregate is valuable for release tracking and dangerous for diagnosis. Two systems with the same 84% correctness can feel completely different: one may fail randomly across low-value questions, while the other fails every permission query from the same customer role.
Build slices around mechanisms of failure:
- lexical lookup versus paraphrase;
- single-hop versus multi-hop evidence;
- stable facts versus temporal policies;
- one source versus conflicting sources;
- answerable versus unanswerable;
- normal access versus permission-sensitive access;
- short documents versus tables, scans, and long structured documents;
- common queries versus rare, high-consequence queries.
Then set gates on important slices, not only on the global mean. A change that raises overall correctness by two points while dropping authorization-sensitive accuracy by ten should not ship.
The BEIR benchmark made a related lesson visible for information retrieval: rankings that look strong in one narrow setting do not necessarily generalize across domains and retrieval tasks. Your evaluation set needs the same heterogeneity your product faces.
How to diagnose one failed query
When a case fails, walk backward from the answer through the recorded trace.
1. Was the question answerable from the authorized corpus? If no, inspect the abstention behavior. If yes, continue.
2. Did the required evidence appear in the retriever’s candidate set? If no, the failure belongs to ingestion, parsing, indexing, query construction, filtering, or retrieval.
3. Did the evidence survive into the final context? If no, inspect reranking, deduplication, chunk expansion, ordering, and token-budget truncation.
4. Did the final context preserve every expected claim? If no, the problem is context construction even if the source document was technically present.
5. Did the generated answer use the available evidence faithfully? If no, the generator or its instructions failed.
6. Did the judge score the case correctly? Evaluation code is software. It has bugs, model drift, ambiguous rubrics, and versioning problems of its own.
This sequence prevents a common waste: changing the prompt when the evidence never reached the prompt.
A minimal evaluation harness
You do not need a large platform to start. You need stage traces, stable cases, versioned configuration, and a comparison against a baseline.
type EvalCase = {
id: string;
question: string;
answerable: boolean;
expectedClaims: string[];
evidenceIds: string[];
slices: string[];
};
type RagTrace = {
query: string;
candidates: Array<{ id: string; score: number }>;
ranked: Array<{ id: string; score: number }>;
context: Array<{ id: string; text: string }>;
answer: string;
citations: string[];
versions: {
corpus: string;
retriever: string;
reranker: string;
prompt: string;
generator: string;
judge: string;
};
};
type EvalResult = {
candidateRecall: number;
contextCoverage: number;
faithfulness: number;
completeness: number;
abstentionCorrect: boolean;
};
Persist the trace before scoring it. A result without the retrieved items, exact final context, and component versions cannot explain a regression.
Run the same cases in three modes:
- Component tests for retrieval and ranking, without generation.
- End-to-end tests for the user-visible answer and latency.
- Paired comparisons between the current baseline and one proposed change.
Change one major variable at a time. If you update the chunker, embedding model, reranker, prompt, and generator together, a better final score teaches you almost nothing. Paired runs give every case the same question and corpus, making the comparison far easier to interpret.
Track cost and latency beside quality. Candidate counts, rerank depth, context length, and judge calls all have operational prices. The best configuration is the one that satisfies the quality gates within the product’s latency and cost budget, not the one that maximizes a leaderboard score without constraints.
Turn production failures into permanent tests
An evaluation suite is not a document completed before launch. It is the memory of the system’s mistakes.
The loop is straightforward:
- Capture traces with privacy and access controls intact.
- Sample failures, low-confidence answers, refusals, and user corrections.
- Have a domain expert label answerability, expected claims, and evidence.
- Add the case to the right diagnostic slices.
- Reproduce the failure offline and fix the responsible layer.
- Run the full suite against the frozen baseline before deployment.
- Keep the case so that failure cannot return silently.
Online feedback belongs in this loop, but a thumbs-up is not a truth label. Users reward tone, speed, and confirmation of what they already believe. Treat feedback as a signal for sampling and review, not as an automatic correctness score.
Monitor production distributions too. Offline quality can remain flat while the share of table questions, new policy versions, or unanswerable requests changes. The system did not necessarily regress; the work arriving at the system changed. Slice volume reveals that shift.
What I would build first
For a new RAG product, my first evaluation milestone is deliberately small:
- 30–50 reviewed cases from real workflows;
- explicit answerable and unanswerable examples;
- stable evidence identifiers and corpus versions;
- candidate retrieval recall at the real reranker cutoff;
- context coverage at the real generation cutoff;
- claim-level faithfulness and expected-claim completeness;
- abstention and over-refusal rates;
- exact traces for every failed case;
- slice-level gates for the two or three highest-risk journeys;
- one frozen baseline used for paired comparisons.
That foundation is more valuable than a dashboard with twenty opaque metrics. Add judge ensembles, synthetic generation, statistical confidence intervals, and online experiments after the traces and labels are trustworthy.
The research agrees on the central decomposition. RAGAS separates context relevance, faithfulness, and answer relevance. ARES evaluates context relevance, answer faithfulness, and answer relevance with calibrated judges. RAGChecker pushes toward finer-grained diagnosis of retriever and generator behavior. The metric names vary. The architectural lesson does not: evaluate the parts if you want to improve the system.
A short opinion
Most RAG teams do not have a model problem first. They have an observability problem.
They can show a polished answer in a demo, but they cannot reconstruct which query ran, which index version answered it, which documents were filtered out, which chunks were reranked, which text reached the model, or which claims the evidence supported. When the answer fails, the team debates prompts because prompts are the only visible component.
The highest-leverage improvement is to make the evidence path inspectable and turn its contracts into tests. Once that exists, model changes become engineering decisions rather than rituals. You can say exactly which slice improved, which layer caused it, what it cost, and what regressed.
That is also the production philosophy behind NEXUS Volume I: retrieval-augmented generation becomes reliable when every layer is measurable, replaceable, and accountable for a clear contract.
References and further reading
- Lewis et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. The paper that formalized the modern RAG architecture around parametric and non-parametric memory.
- Karpukhin et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. The canonical dense-retrieval baseline behind many early RAG systems.
- Petroni et al. (2021). KILT: a Benchmark for Knowledge Intensive Language Tasks. A benchmark that couples answers with provenance across several knowledge-intensive tasks.
- Thakur et al. (2021). BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. Evidence that retrieval performance must be tested across varied datasets and tasks.
- Es et al. (2024). RAGAS: Automated Evaluation of Retrieval Augmented Generation. Reference-free metrics separated across retrieval and generation dimensions.
- Saad-Falcon et al. (2024). ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. Automated judges anchored by human annotations and prediction-powered inference.
- Ru et al. (2024). RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation. Claim-level metrics designed to reveal retriever and generator failure patterns.
- Liu et al. (2024). Lost in the Middle: How Language Models Use Long Contexts. Why the presence and position of evidence in a long context do not guarantee that a model will use it.