RAG Evaluation: How to Measure Retrieval Quality and Answer Accuracy
A RAG system can return polished, confident answers and still be performing badly. The retriever could be missing the right document, giving irrelevant chunks higher ranks than useful evidence, or passing incomplete context to the model. Then the model can provide an answer that is plausible but not fully supported.
That’s why RAG evaluation should measure more than just the final response. A production evaluation plan differentiates retrieval quality from answer quality, tests representative business questions before launch, establishes action thresholds, and continues measuring after deployment.
This guide explains how to measure retrieval quality and answer accuracy, how to construct a useful test set, when human review still matters, and how to turn evaluation results into release decisions instead of a dashboard nobody acts on.
What RAG Evaluation Measures in Practice
RAG evaluation measures how well a system finds the right evidence and how reliably the model uses that evidence to answer the user.
A useful evaluation splits the pipeline into at least two levels. Retrieval evaluation asks whether the system has found the evidence it should have found. Answer evaluation asks whether the answer is relevant, complete, grounded in that evidence, factually correct if there is a reference answer, and whether it has correct citations.
That separation matters because a weak answer can have different causes. Changing the prompt will not fix the root problem if the correct policy section was never retrieved. If the right evidence was retrieved but the model ignored it or hallucinated, the retrieval layer may be fine and the generation layer needs work.
Microsoft Foundry now exposes RAG-specific evaluators for dimensions such as groundedness, response completeness, retrieval quality, and relevance. The RAGAS research framework similarly treats retrieval and generation as separate evaluation problems rather than one combined score. Microsoft’s RAG evaluator guidance is a useful reference for teams defining their own evaluation set.
Original planning framework: separate retrieval, answer quality, release rules, and production monitoring.
Start With Retrieval Quality, Not the Final Answer
The retriever controls what evidence the model observes. Precision and recall are a practical starting point when questions are tied to a known relevant set of documents or passages.
Precision@K
Precision@K asks how many of the top K retrieved results were actually relevant. If the system returns four chunks and two are useful, Precision@4 is 2 divided by 4, or 0.50. Low precision means the model is receiving noise that can distract it, waste context space, and increase the chance of an unsupported answer.
Recall@K
Recall@K measures the fraction of known relevant evidence retrieved in the top K results. So if the labelled test set says we need 2 passages, and those 2 passages appear in the top 4 results, Recall @ 4 is 2/2, or 1.00. Recall of 0.50 if only one appears.
When the corpus does not have a labelled set of relevant passages, it is harder to calculate recall. That is a reason to create a curated test set for production work, not to just throw the metric away. The team does not have to annotate the entire knowledge base. For one to compare the changes in retrieval and to identify regressions, a sufficient number of representative questions is needed.
Illustrative worked example
Assume an internal policy assistant receives the question, “What approvals are required for travel over $5,000?” The test owner has labelled two source passages as required evidence: the travel policy section and the approval matrix.
| Retrieved rank | Result | Relevant? | Why it matters |
| 1 | Travel Policy – approval rules | Yes | Contains the approval threshold and policy condition |
| 2 | Expense FAQ – meal limits | No | Related to travel, but not to approval authority |
| 3 | Approval Matrix – finance authority | Yes | Contains the second required source |
| 4 | Travel booking instructions | No | Operationally related, but does not answer the approval question |
In this example, Precision@4 is 0.50 and Recall@4 is 1.00. The retriever found all required evidence, but half of the available context is noise. That suggests a ranking or filtering problem rather than a missing-source problem. The next test should compare changes such as metadata filtering, hybrid search, reranking, query rewriting, or different chunking against the same labelled question set.
Measure the Answer on More Than “Looks Good”
Once retrieval is measured separately, evaluate what the model does with the retrieved context. A useful scorecard usually needs several dimensions because a single quality score can hide different failure modes.
| Measure | What it asks | Typical evidence |
| Groundedness | Is each material claim supported by the retrieved context? | Claim-to-source comparison or evaluator score |
| Response completeness | Did the answer include the important information required by the question? | Reference answer, rubric, or required facts |
| Answer relevance | Did the response directly answer the user instead of drifting into related material? | Human rubric or model-based evaluator |
| Correctness | When ground truth exists, is the answer factually correct? | Gold answer, source record, or structured truth set |
| Citation quality | Do citations point to the right source and support the nearby claim? | Citation-to-source validation |
| Safety and authorization | Did the answer avoid restricted or disallowed information? | Negative access tests and security test cases |
Do not collapse all of those measures into one average unless the weighting reflects business risk. A 95 percent average can still hide a serious security failure, a consistently missing citation, or poor recall on the highest-value use case.
Build a Test Set Before You Tune the System
A good RAG test set should mirror the type of traffic that the system will actually encounter. It should have simple questions, but the most useful cases are often those that expose ambiguity, source conflicts, permission boundaries, stale information, and lack of evidence.
- Common questions that represent the expected production workload.
- Questions that require more than one document or more than one section of a document.
- Ambiguous wording, acronyms, synonyms, and internal terminology.
- Questions where the correct behaviour is to say that the available sources do not support an answer.
- Questions involving recently changed policies or documents with multiple versions.
- Permission-restricted questions that a particular test user must not be able to answer.
- High-risk questions where a wrong or incomplete response would create material operational, legal, financial, or customer impact.
Capture the expected evidence, expected answer characteristics, user identity or permission context and business severity of failure where possible. This transforms the testset from a spreadsheet of sample prompts to an engineering asset.
The value of the evaluation set is directly dependent on the quality of the source corpus. Before tuning retrieval repeatedly, fix the data layer if the test reveals stale documents, broken parsing, poor metadata, or inconsistent versioning. Our guide to RAG data preparation (coming soon) covers the source, metadata, chunking, permission, and update decisions that sit underneath evaluation.
Use Human Review and Model-Based Evaluation for Different Jobs
LLM-based judges are a way to make regression testing more practical by reducing the cost of repeated evaluation, but they should not be considered unquestioned ground truth. Their prompts, rubrics, choice of model, and reference context can all influence the score.
Use automated evaluators for frequent comparisons and wide regression coverage. Maintain human review for high-impact decisions, ambiguous cases, evaluator calibration, and periodic spot checks. A good operating pattern is to compare model-based scores against human ratings on a small labelled sample prior to deploying the evaluator at scale.
The RAGAS paper is useful here because it formalizes reference-free evaluation dimensions, but any framework still needs to be tested against the organization’s own use case. The right metric is the one that predicts whether the system is safe and useful in production.
Turn Metrics Into Release Thresholds
Evaluation becomes operational when every important measure has an owner and an action rule. Avoid universal numbers copied from another system. A customer-support assistant and a policy assistant may need very different tolerances.
| Area | Example release rule | Owner |
| Restricted information | No unauthorized retrieval or disclosure in the defined security test suite | Security and application owner |
| Critical retrieval | Required evidence must be retrieved for every labelled critical question before release | RAG engineering owner |
| Citations | Material factual claims in high-risk workflows must resolve to supporting sources | Product or compliance owner |
| Answer completeness | Known required facts must be present for the critical reference set | Business process owner |
| Regression | A proposed change cannot materially reduce an agreed core metric without review | Technical owner |
Thresholds should trigger a response. A retrieval regression might block deployment. A drift signal might trigger a fresh test run after a major corpus update. A citation failure might send the affected query to human review. The measurement is useful because it changes what the team does next.
| NEXT STEP If your team is defining how to test a RAG system before production, Arcadion’s RAG Systems team can help review the retrieval architecture, evaluation plan, test set, and production controls. |
Monitor Production RAG as the Corpus and Users Change
Pre-launch score is not a permanent quality assurance. Documents change, permissions change, new questions from users, model or retrieval updates can change behaviour.
Production monitoring should be a combination of operating metrics and quality assessment by sampling. Track retrieval failures, zero result queries, latency, token and inference cost, source freshness, fallback rates, citation failures, user feedback, and any queries escalated for human handling. Pair these signals with periodic runs over the labelled evaluation set.
Explore by layer when a metric changes. Has the corpus changed? Did something change in a parser? Has chunking/embeddings changed? Has the retrieval query or ranking logic changed? Was the prompt or model changed? At the layer level, this diagnosis can be made much more quickly.
A Practical RAG Evaluation Checklist
- Define the business outcome and the failures that matter most.
- Create a representative question set with expected evidence where feasible.
- Measure retrieval separately from final answer quality.
- Use precision and recall where the relevant set can be labelled.
- Assess groundedness, completeness, relevance, correctness, citations, and authorization behaviour.
- Calibrate automated evaluators against human review.
- Set release thresholds and name the owner of each threshold.
- Re-run the same tests after changes to data, retrieval, prompts, models, or permissions.
- Monitor production queries and feed new failure cases back into the test set.
RAG Evaluation Should Make Release Decisions Easier
The goal of RAG evaluation is not to build the biggest possible dashboard. The aim is to surface the failure modes of the system before users encounter them and to give the team a repeatable way to compare changes.
A strong program tests the retrieval layer, the answer layer, and the security boundary individually. It employs representative questions, explicit thresholds, and production feedback. This makes “is this ready?” a question of evidence, not a subjective review of a handful of compelling demos.
Planning a production RAG system?
Explore Arcadion’s RAG Systems capabilities to see how our team supports data ingestion, retrieval configuration, prompt engineering, testing, tuning, and governance.
