Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Check retrieval and the generated answer as two separate stages. First, test whether relevant, answer-bearing passages appear high enough in the results; then check whether the answer is accurate, complete, grounded in those passages, and properly cited. A fluent answer or a high aggregate score alone cannot prove the system found the right documents.

Define what “the right documents” means for your question

In a retrieval-augmented generation (RAG) system, a retrieval system or knowledge base identifies information relevant to a query and supplies it to the AI model as context. That is how NIST defines RAG in its glossary.

Relevance is task-specific. A passage can share the query’s topic but still omit the fact needed to answer it. Before testing, pair each representative query with the documents or passages that contain useful evidence, and decide what counts as sufficient evidence for that question. Microsoft’s search evaluation guidance recommends preparing test queries and the text in test documents that address them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Test whether retrieval found and ranked the evidence

  1. Build a representative test set. Use questions real users ask and identify the relevant documents or passages in advance. Include answerable questions as well as questions the corpus cannot answer; Microsoft recommends positive and negative examples.
  2. Inspect retrieved passages before judging the answer. For each query, check whether the answer-bearing evidence appears in the results and whether irrelevant material crowds it out. This isolates retrieval performance from what the model does with its context.
  3. Measure relevance, coverage, and rank separately. Precision at K is the share of the top K results judged relevant. Recall at K is the share of all relevant items found among those results. Mean Reciprocal Rank (MRR) reflects how high the first relevant result appears. AWS also distinguishes context relevance from context coverage, which evaluates retrieved context in relation to ground-truth texts.
  4. Check for missing evidence, not just relevant snippets. A short passage can be relevant yet insufficient. Compare the retrieved context with the known answer-bearing material and note missing documents or facts.
  5. Evaluate the answer as a second stage. Check correctness and completeness, whether claims are supported by retrieved context, and whether citations point to passages that actually support the claims.
  6. Record failures per query. Averages can hide a serious miss on a consequential question. Keep the failed queries and their evidence visible when reviewing results.
  7. Compare changes on the same test cases. When you change indexing, retrieval, or ranking settings, run the same queries and relevance judgments again. Otherwise, score differences may reflect a different test set rather than a better system.

Choose metrics that answer distinct questions

Evaluation question Useful measure What it tells you
Are the top results pertinent? Precision at K or context relevance Whether returned passages are relevant and how much irrelevant material appears.
Did retrieval find enough evidence? Recall at K or context coverage Whether relevant items or answer-bearing information are missing from the retrieved set.
Is useful evidence near the top? MRR or a rank-sensitive measure such as nDCG How highly useful results rank. Microsoft describes MRR; NIST’s TREC evaluation reports nDCG and recall.
Is the answer accurate and responsive? Correctness and completeness Whether the response is right and addresses the question.
Are answer claims grounded in the retrieved context? Faithfulness or groundedness Whether claims are supported by the supplied evidence.
Do citations support claims, and are claims cited? Citation precision and citation coverage Whether cited passages are correct and how well the response is supported by citations.

K is the cutoff used for a top-results metric; choose it to fit the task rather than treating one cutoff as universal. Scores also depend on the test queries and relevance judgments, so do not assume scores from different test sets are directly comparable. Microsoft advises examining positive and negative query results separately when reviewing aggregate behavior.

Use failures to locate the problem

  • Top results are mostly irrelevant: investigate retrieval quality and ranking.
  • Some relevant passages appear, but key facts or documents are absent: investigate coverage.
  • The evidence was retrieved, but the answer misstates or ignores it: investigate generation correctness and faithfulness.
  • The answer may be right, but citations do not support its claims: investigate citation precision and coverage.
  • An unanswerable query produces confident, irrelevant material: check how the system handles questions for which the corpus has no answer.

This separation matters: finding a relevant passage does not prove the answer used it correctly, and a plausible answer does not establish that retrieval found the necessary evidence.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What published evaluations can—and cannot—tell you

NIST’s July 18, 2025 study, updated September 18, 2025, reports that in the TREC 2024 RAG Track, rankings based on automatically generated UMBRELA relevance assessments correlated highly with rankings based on fully manual assessments for nDCG@20, nDCG@100, and Recall@100 across 77 runs from 19 teams. In that study, LLM assistance did not appear to increase correlation with fully manual assessments. This is evidence about run-level effectiveness in that benchmark, not a guarantee that automated judgments will be reliable for another corpus or for an individual system decision. See NIST’s study.

NIST’s overview of the TREC 2025 RAG Track describes four tasks: retrieval, augmented generation, retrieval-augmented generation, and relevance judgment. Its support evaluation uses weighted precision to measure correct passage citations and weighted recall to measure how many answer sentences are supported by passage citations. These are evaluation designs, not universal pass thresholds.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The cited guidance defines useful measures and reports benchmark findings, but does not establish a score that proves a system always finds the right documents. Use metrics to compare performance on your task, and inspect consequential misses directly.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.