Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To tell whether RAG retrieval is broken, evaluate what documents the retriever returns separately from how the model uses those documents. Check retrieval against relevance judgments when you have them; then evaluate the answer for groundedness, relevance, and completeness. Finally, inspect individual traces so a weak answer can be tied to the retrieval step—or to what happened after retrieval. No single score or pass threshold proves a RAG system is reliable.

What does “broken retrieval” mean?

Retrieval is the upstream process of finding useful evidence for a query. It can fail by missing relevant documents, ranking them too low to make the context window, or returning so much irrelevant text that useful evidence is crowded out. Those are different problems from a model that receives the right evidence but ignores it, misstates it, or leaves out an important part of the answer.

Microsoft Foundry distinguishes process evaluation from system evaluation: retrieval metrics assess the search process, while answer-level evaluation assesses the resulting response. That distinction is practical, not merely terminological. A correct-looking answer does not show that the retriever found the right evidence, and a poor answer does not by itself show that retrieval was the cause. Microsoft Foundry evaluation approach

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose retrieval metrics based on your evidence

When you have relevance labels

If people have identified which documents are relevant to each test query, compare the retrieved documents with those judgments. Microsoft Foundry’s labeled document retrieval evaluator includes metrics such as Fidelity, NDCG, XDCG, Max Relevance, and Holes. NDCG helps assess whether relevant results are ranked well; Holes can flag gaps in the relevance judgments themselves. An evaluation set with missing judgments can make a sound result look irrelevant—or an irrelevant result appear acceptable—so inspect the labels as well as the scores. Microsoft Foundry RAG evaluation

Look at the returned documents and their rank, especially within the number of results your application actually passes to the model. A useful document ranked below that cutoff is effectively absent from the answer-generation context. When query-level labels are available, this comparison is the more direct way to measure whether known relevant documents were found and ranked.

When you do not have labels

A model-judged context or retrieval relevance measure can provide an initial signal: does the retrieved text appear useful for the query? It is useful for triage, but it is not equivalent to comparing search results with human-labeled relevant documents. The score depends on the evaluator, so review representative examples and add human judgments where the consequences of a wrong decision justify the effort. Microsoft Foundry RAG evaluation

Evaluate the answer separately from retrieval

Once you know what context reached the model, assess the response along distinct dimensions. Microsoft Foundry describes groundedness, relevance, and completeness evaluators; Microsoft’s architecture guidance recommends combining evaluation dimensions rather than relying on one measure. Microsoft Foundry RAG evaluation Microsoft Azure architecture guidance on RAG metrics

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Groundedness: Are the answer’s claims supported by the retrieved context? A groundedness or faithfulness check can catch unsupported statements, but it does not establish that retrieval found all the evidence needed.
  • Answer relevance: Does the response address the user’s actual question? A response can be supported by context and still answer the wrong question.
  • Completeness: Does the answer include the expected critical information? This helps catch omissions that a groundedness check alone may miss.

Ragas also includes faithfulness among its RAG metrics, and some of its metrics use LLM calls. That makes evaluator choice and interpretation part of the measurement, rather than a reason to treat an automated score as ground truth. Ragas metric reference

Build an evaluation set that reflects real queries

Start with realistic questions and the evidence or answer characteristics that matter for each one. Include ordinary queries as well as difficult forms: Google Cloud recommends varied golden questions, including simple, complex, multi-part, and misspelled examples. Refresh the set as real usage and application requirements change; a fixed set can stop representing the workload. Google Cloud guidance on evaluating search quality

For each case, preserve enough information to diagnose the result: the query, expected relevant evidence or judgments, retrieved documents, generated answer, and evaluator outputs. Logging retrieval intermediates makes it possible to trace an answer failure to the point where it arose, rather than infer the cause from the final text alone. Databricks guidance on evaluation and monitoring

Run a controlled comparison

  1. Define the failure you want to detect. Specify whether the concern is missing evidence, irrelevant context, unsupported claims, poor relevance, or omissions. Choose measures that match those failure types.
  2. Establish a baseline. Run the representative query set through the current system and record retrieved documents and answers alongside the scores.
  3. Change a retrieval choice deliberately. Candidate dimensions include the search algorithm, top-k, and chunk size. Where practical, change one at a time so a difference is easier to attribute.
  4. Compare both stages. Check relevant evidence found, ranking, and context noise, then check groundedness, answer relevance, and completeness. Consider latency, cost, or implementation complexity when they matter to the application.
  5. Inspect failures and update the set. Review examples with people, determine whether the cause is retrieval, answer generation, or the evaluation set, and add newly important cases to future runs.

Microsoft Foundry describes parameter sweeps across retrieval algorithms, top-k, and chunk sizes; Google Cloud recommends iterative runs against a baseline. Keep the same evaluation set across configurations so the comparison is meaningful. Microsoft Foundry RAG evaluation Google Cloud guidance on evaluating search quality

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Interpret scores as signals, not certification

Scores depend on the evaluator and the workload. Microsoft Foundry’s listed evaluators return scores from 1 to 5 and document a default pass threshold of 3. That is an implementation default in Microsoft’s evaluator documentation, not a universal definition of acceptable RAG quality or an independently validated benchmark. Microsoft Foundry RAG evaluation

There is no universal acceptable threshold established here for recall@k, NDCG, faithfulness, or completeness. Set targets according to the application’s needs and the cost of failure, and calibrate them against reviewed examples. Model responses can be nondeterministic, so interpret aggregate scores alongside individual outputs and human review. Microsoft Azure architecture guidance on RAG metrics Google Cloud guidance on evaluating search quality

Use the failure pattern to choose the fix

  • Relevant evidence is absent: investigate retrieval coverage and the search configuration; answer-generation changes cannot use documents the model never received.
  • Relevant evidence is present but ranked too low or crowded out: inspect ranking, top-k, chunk size, and context noise.
  • Evidence is present, but the answer makes unsupported claims: investigate how the model uses context and assess groundedness against the retrieved text.
  • The answer is supported but misses a required point: check completeness against expected information; groundedness alone does not test whether all necessary content was included.
  • Scores and examples disagree: review the evaluator, relevance judgments, and representative traces before changing the system based on the score alone.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.