Skip to main content

How to Evaluate a RAG System — Retrieval Quality vs Answer Quality

Calculating read time…

A traditional application often has a relatively direct relationship between input and output. A RAG system is different. Before the final answer is produced, the system usually has to find relevant information, rank that information, construct the context, and then ask a language model to generate a response from it.


That creates an important engineering challenge: the final answer does not tell you exactly where the failure occurred. A wrong answer could be caused by poor chunking, weak retrieval, incorrect ranking, noisy context, a generation problem, or unsupported reasoning by the language model.

This is why evaluating only the final answer is not enough. A production-quality RAG evaluation process should examine the system at multiple levels.

📍
The core idea

First ask: Did RAG retrieve the right evidence?
Then ask: Did the model use that evidence correctly?

The Two Layers of RAG Evaluation

The cleanest way to understand RAG evaluation is to separate the problem into two major layers: retrieval quality and answer quality.

Layer 01

Retrieval Quality

Did the system find the information required to answer the question?

This includes retrieval recall, precision, ranking quality, context relevance, metadata filtering, hybrid search, reranking, and top-K selection.

Layer 02

Answer Quality

Did the model produce a useful, correct, and evidence-supported answer?

This includes faithfulness, relevance, correctness, completeness, citation support, and appropriate handling of uncertainty.


One Question Can Reveal the Entire Problem

Consider an employee asking:

“What is the maximum amount I can claim for international business travel?”

Suppose the correct policy exists somewhere inside the company's travel documentation. There are several ways the RAG system can fail.

What happened? Likely failure Where to investigate
The correct policy was never retrieved. Retrieval failure Chunking, embeddings, search, filters, ranking
The correct policy was retrieved, but the model gave the wrong number. Generation / grounding failure Prompt, context construction, model behavior
The answer contains information that is not supported by the retrieved policy. Faithfulness failure Claim grounding and generation
The answer discusses the policy but never answers the actual question. Answer relevance failure Prompt and response generation

Part 1 — Evaluating Retrieval Quality

Retrieval is the foundation of a RAG system. If the required evidence never reaches the model, the generation layer has very little chance of producing a reliably grounded answer.

Retrieval evaluation therefore focuses on questions such as: Did we find the relevant information? Did we retrieve enough of it? Did we rank it highly enough? And how much irrelevant material did we bring along?

Recall@K — Did We Find the Important Evidence?

Recall@K measures whether the relevant information appears within the first K retrieved results. In simple terms, it answers: “How much of the information we needed did our search actually find?”

Recall@K = Relevant evidence retrieved in top K ÷ Total relevant evidence

Imagine that a question requires four relevant chunks and your top-10 retrieval results contain three of them. The system found most of the required evidence, but one relevant piece is still missing.

Recall becomes especially important for questions where the answer depends on multiple pieces of information. Missing one critical chunk can make an otherwise strong RAG system produce an incomplete answer.

Precision@K — How Much of What We Retrieved Is Useful?

Precision looks at the other side of the problem. Instead of asking what we missed, it asks how much of what we retrieved is actually relevant.

Precision@K = Relevant results in top K ÷ K

Suppose the system retrieves ten chunks, but only three are useful. The remaining seven chunks add noise. Depending on the application, that extra context can increase token usage, latency, and the chance that important evidence gets buried among irrelevant information.

MRR — How Quickly Do We Find a Relevant Result?

Mean Reciprocal Rank, or MRR, focuses on the position of the first relevant result. If the first useful result appears at rank 1, its reciprocal rank is 1. If it appears at rank 4, the value becomes 1/4.

MRR is particularly useful when finding one strong result quickly is important. It is less expressive when the answer requires several different pieces of evidence.

NDCG — Are the Results Ranked Properly?

Not every relevant chunk has the same value. One passage might directly answer the question while another merely provides background information.

NDCG, or Normalized Discounted Cumulative Gain, evaluates ranking while considering different levels of relevance. It is useful when your evaluation dataset can distinguish between highly relevant, somewhat relevant, and irrelevant results.

Retrieval Metrics at a Glance


Metric Main question What it helps diagnose
Recall@K Did we retrieve the important evidence? Missing information
Precision@K How much retrieved content is relevant? Retrieval noise
MRR How high is the first relevant result? First-result ranking
NDCG@K Are highly relevant results ranked well? Ranking quality

Part 2 — Evaluating Answer Quality

Finding the right information is only half of the RAG problem. The model still needs to interpret the retrieved context and produce an answer that is grounded, relevant, correct, and complete.

Faithfulness — Is the Answer Supported by the Evidence?

Faithfulness asks whether the claims made by the generated answer can be supported by the retrieved context.

This is one of the most important concepts in RAG because a language model can generate a fluent statement that was never actually present in the retrieved evidence.

Example

The retrieved policy says: “Employees receive 20 days of annual leave.”

If the model adds, “Unused days can be carried forward indefinitely,” that additional statement needs its own supporting evidence.

A response can therefore be fluent, relevant, and even partially correct while still having a faithfulness problem.

Answer Relevance — Did We Actually Answer the Question?

An answer can contain accurate information and still fail to satisfy the user's request.

For example, a user asks, “What is the maximum reimbursement amount?” The model responds with three paragraphs describing the reimbursement policy but never gives the maximum amount. The information may be related and factually correct, but the answer is not sufficiently aligned with the question.

Correctness — Is the Answer Actually Right?

Correctness compares the generated response against a trusted reference answer or an appropriately defined expected outcome.

Exact string matching is often too restrictive for natural language. Two answers can use completely different wording while expressing the same fact. For open-ended questions, human evaluation or carefully designed model-based evaluation can therefore be more useful than simple text matching.

Completeness — Did the Answer Cover Everything Required?

Some questions require several pieces of information. An answer may correctly provide two of them while silently missing the third.

This is why correctness and completeness should not automatically be treated as the same thing. An answer can contain only correct statements and still fail to fully satisfy a multi-part question.

The Most Useful Diagnostic Matrix

This simple matrix is one of the most practical ways to reason about RAG failures.

Retrieval Answer What it suggests
Good Good The pipeline is behaving as intended.
Poor Poor Investigate retrieval first.
Good Poor The evidence exists; investigate generation and context usage.
Poor Good-looking The answer may be correct for reasons unrelated to the retrieved evidence.

Follow the Question Through the RAG Pipeline

A useful evaluation mindset is to trace a question through every major stage instead of jumping directly to the final answer.

User Question
→
Retrieval
→
Context
→
LLM
→
Answer
Evaluate each transition instead of treating the final answer as a black box.

Build the Evaluation Dataset Before Chasing Metrics

Metrics are only meaningful when the evaluation questions represent the actual job your RAG system is expected to perform.

A useful evaluation dataset should contain realistic questions rather than a collection of artificially simple examples. It should cover straightforward questions, difficult questions, multi-part questions, questions requiring multiple documents, ambiguous questions, and questions for which the knowledge base contains no reliable answer.

A strong evaluation dataset asks more than “What is the correct answer?”

It should also capture which evidence supports that answer and whether the system retrieved that evidence.

What Should One Evaluation Record Contain?

Question — What did the user ask?

Reference answer — What should a correct answer contain?

Reference evidence — Which source material supports the answer?

Retrieved chunks — What did the RAG system actually retrieve?

Generated answer — What did the LLM produce?

Evaluation results — How did retrieval and generation perform?

Your Test Set Should Include Questions the System Cannot Answer

This is an often-overlooked part of RAG evaluation.

If every test question has a clean answer inside the knowledge base, you are mainly testing whether the system can retrieve and generate an answer. You are not testing whether it can recognize the boundary of its available knowledge.

A serious evaluation set should therefore include unsupported questions, outdated information, ambiguous requests, conflicting information, and questions outside the intended scope of the application.

These cases help evaluate whether the system can respond appropriately when the evidence is insufficient instead of confidently inventing an answer.

Where Does LLM-as-a-Judge Fit?

Natural-language answers are difficult to evaluate using only exact-match metrics. An LLM can therefore be used as an evaluator to assess properties such as relevance, faithfulness, completeness, or similarity to a reference answer.

This can dramatically increase evaluation scale, but it introduces another model into the evaluation process. The evaluator can itself make mistakes or interpret the scoring criteria inconsistently.

Use an LLM judge as an evaluator—not as an unquestionable source of truth. Validate automated judgments against human-reviewed examples, especially for high-impact use cases.

Why Human Evaluation Still Matters

Automated evaluation is excellent for scale, regression testing, and rapid experimentation. Human evaluation remains important for validating whether the evaluation criteria actually match what users consider useful and correct.

A practical approach is to use human-reviewed examples to calibrate automated evaluation, investigate difficult failures, and periodically verify that automated scores continue to reflect real quality.

A Practical RAG Evaluation Scorecard

Dimension Typical evaluation Core question
Retrieval recall Recall@K / Context Recall Did we find the required evidence?
Retrieval precision Precision@K / Context Precision How much retrieved context is useful?
Ranking MRR / NDCG Are useful results ranked highly?
Faithfulness Claim-to-context evaluation Is the answer supported by evidence?
Answer relevance Automated + human review Did we answer the question?
Correctness Reference / human evaluation Is the answer actually correct?
Completeness Requirement / claim coverage Did we cover everything required?

How to Debug a Bad RAG Answer

When a RAG answer fails, the first instinct is often to change the prompt. That can be useful, but it should not be the first diagnostic step.

Instead, trace the failure backward through the pipeline.

Debugging sequence
Did we retrieve the right document?
↓
Did we retrieve the right chunk?
↓
Was the relevant chunk ranked highly enough?
↓
Was the correct context passed to the model?
↓
Did the model use that evidence?
↓
Is every important claim supported?
↓
Did the answer actually satisfy the question?

What Should You Change When Evaluation Is Poor?

Observed problem Areas worth investigating
Relevant evidence is frequently missing Chunking, embeddings, query formulation, search strategy, metadata filters
Relevant evidence is retrieved but ranked poorly Ranking, hybrid search, reranking, top-K configuration
Relevant context exists but answer is unsupported Prompt, context construction, generation behavior
Answer is correct but incomplete Question decomposition, evidence coverage, answer instructions
Answer is fluent but irrelevant Prompting, response structure, answer relevance evaluation

Offline Evaluation vs Production Monitoring

A test dataset tells you how your system behaves on known examples. Production traffic exposes questions, documents, and failure patterns that you may not have anticipated.

Offline Evaluation

Controlled test set, repeatable experiments, regression testing, model comparison, retrieval tuning, and prompt evaluation.

Production Monitoring

Real user questions, unexpected queries, retrieval failures, latency, cost, user feedback, and new failure patterns.

A mature RAG system creates a feedback loop between the two: production failures become new evaluation cases.

Don't Optimize for One Number

One of the easiest mistakes in RAG evaluation is to treat a single metric as the definition of system quality.

Increasing the number of retrieved chunks may improve recall because the system sees more information. At the same time, precision may decrease, context may become noisier, token consumption may increase, and latency may rise.

Likewise, an answer may score well for relevance while still containing unsupported claims. A system may retrieve excellent evidence while failing to communicate it clearly to the user.

Think in dimensions, not one magic score.

Retrieval tells you whether the evidence was found. Faithfulness tells you whether the answer stayed grounded. Relevance tells you whether the response addressed the question. Correctness and completeness tell you whether the user actually received the required information.

A Practical RAG Evaluation Workflow

1. Define the scope. Decide what the system should answer and where its knowledge boundary begins and ends.

2. Build representative questions. Include real user tasks, difficult cases, multi-part questions, and unsupported questions.

3. Identify supporting evidence. Record which source material actually supports each expected answer.

4. Evaluate retrieval. Measure whether the relevant evidence appears in the retrieved results and whether it is ranked appropriately.

5. Evaluate the answer. Measure grounding, relevance, correctness, and completeness.

6. Inspect failures. Determine whether each failure came from retrieval, context construction, generation, or the evaluation itself.

7. Make controlled changes. Change one important variable at a time whenever possible.

8. Re-run the complete evaluation. Look for improvements and regressions across the entire scorecard.

The Mental Model to Remember

A RAG system must answer two questions.
Did we find the right evidence?
Retrieval Quality

Did we use that evidence correctly?
Answer Quality

Final Takeaway

Evaluating RAG is not simply a matter of checking whether the final response “looks good.” A RAG system is a chain of retrieval and generation decisions, and each part can fail independently.

If the correct evidence was never retrieved, the problem belongs primarily to the retrieval layer. If the evidence was retrieved correctly but the model produced an unsupported or incomplete response, the problem shifts toward context usage and generation.

That is why strong RAG evaluation follows the architecture itself: measure retrieval, measure grounding, measure answer quality, inspect failures, and then evaluate the complete user experience.

If you remember only two questions

1. Did we retrieve the right evidence?
2. Did the answer stay faithful to that evidence?

Once you start looking at RAG this way, evaluation stops being a collection of scores and becomes something much more useful: a method for understanding and improving the system.

Comments