Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To know whether retrieval is good, test it against representative questions with judgments about which passages are relevant. Check retrieval separately from the generated answer: a fluent response does not prove that the system found the right evidence. Then inspect whether the answer is grounded, relevant, complete, and correct for the task.
What does “good retrieval” mean?
A retriever is doing its job when it brings the evidence needed for a question into the material supplied to the language model, and ranks useful evidence where it can be used. That is different from producing a good answer. A model may write convincingly despite poor retrieval, or fail to answer well even when the right passages were found.
Evaluate the retrieval stage and the response as separate parts of the system. For retrieval, ask whether relevant passages were found and ranked usefully. For the response, ask whether its claims are supported by those passages and whether it answers the question fully and correctly.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Which dimensions should you measure?
Retrieved documents and their ranking
Use representative queries and explicit judgments about which documents or passages are relevant. Check both whether the relevant evidence appears in the retrieved set and whether it appears high enough in the ranking to be useful. The right emphasis depends on your workload; a general benchmark may not reflect the kinds of queries your application receives.
#1 Best Overall
Microsoft Foundry documents a relevance-labeled document retrieval evaluator with measures including Fidelity, NDCG, XDCG, Max Relevance, and Holes. These are names used by that implementation, not a universal set exposed by every framework. Choose measures based on the questions your evaluation needs to answer, rather than treating any one score as a complete verdict.
Context passed to the model
Context relevance asks whether the retrieved material is focused on the query. Context recall asks whether the context contains the evidence needed to answer; in the cited Microsoft metric description, an annotated answer serves as a proxy. These dimensions reveal different problems: irrelevant passages add noise, while missing evidence can make a correct answer impossible.
Rank #2
The generated response
Faithfulness or groundedness checks whether answer claims are supported by the supplied context. Answer relevance checks whether the response addresses the query; it does not by itself establish factual correctness. Depending on the task, assess completeness, utilization of evidence, relevance, and correctness as distinct dimensions.
Free tools Windows power users keep installed
One-click scans. No signup required.
RAGAS framed automated RAG evaluation around faithfulness, answer relevance, and context relevance in its 2024 paper. Its current metric documentation also lists faithfulness, answer accuracy, and context relevance. Names and implementations can vary by tool and version, so check the documentation for the version you plan to use before relying on a particular metric.
Rank #3
How do you build a repeatable evaluation?
- Choose representative cases. Include common questions, ambiguous wording, multi-part questions, questions requiring evidence from several passages, and questions the index cannot answer. There is no universal test-set size established here; prioritize cases that reflect how your users actually ask for information.
- Label relevant evidence. For each query, record which documents or passages count as relevant and apply the same relevance scheme consistently. If detailed judgments are costly, begin with a smaller, carefully reviewed set. Automated methods can help expand or triage labels, but validate those labels before using them to draw conclusions.
- Log each run. Keep the query, retrieved passages and their ranks, the context sent to the model, the generated answer, and the configuration or version being tested. Without those details, a poor answer is difficult to attribute to retrieval, context construction, or generation.
- Score retrieval on its own. Use relevance and ranking measures suited to the application, then inspect missed relevant passages and noisy high-ranked results. Treat a framework’s evaluator as one implementation of this task, not as a universal standard.
- Score the response separately. Assess groundedness, relevance, completeness, and correctness according to what the application must do. If the answer is grounded but incomplete or incorrect, inspect whether the evidence was missing, insufficient, or misinterpreted.
- Compare configurations on the same cases. Where feasible, change one meaningful part of the pipeline at a time and rerun the fixed set. Keep results by dimension rather than hiding tradeoffs in one blended score. Model responses can be nondeterministic, so repeat model-based evaluations when variability could affect the decision.
- Review failures by category. Look at missed relevant passages, irrelevant high-ranked passages, unsupported claims, unanswered questions, and differences across query types. An average can conceal a serious failure in a segment of the workload.
How can you tell which part of the pipeline failed?
| Observed result | Likely area to investigate | What to inspect |
|---|---|---|
| Relevant passages do not appear in the retrieved set | Indexing or retrieval | Whether the needed evidence is represented in the index and whether retrieval finds it for the query. |
| Relevant passages appear, but are buried below less useful results | Retrieval ranking | The ordering of relevant and irrelevant passages and the measures used to assess ranking. |
| The retrieved set contains useful evidence, but the model receives irrelevant or insufficient context | Context construction or prompting | Which retrieved passages are passed to the model and whether they provide enough evidence for the question. |
| The answer makes claims unsupported by the supplied passages | Grounding or generation | The unsupported claims alongside the context the model received. |
| The answer is supported by context but misses part of the task or gives an incorrect result | Evidence sufficiency, interpretation, or task-specific response quality | Whether the context contains all required evidence and whether the response uses it correctly and completely. |
Use the retrieved passages and the answer together to diagnose a failure. A retrieval score alone cannot establish that the end-to-end system works well, and a weak answer is not automatically a generator problem if the needed source material never reached the model.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How should you compare results?
Keep a per-dimension view of the tradeoffs. A configuration may improve coverage while adding irrelevant material, or produce better-grounded answers without making them more complete. Track:
- Coverage: whether relevant passages appear in the returned context.
- Ranking: whether the most useful passages are near the top.
- Noise: how much irrelevant material is included.
- Answer grounding: whether claims are supported by retrieved evidence.
- Answer usefulness: whether the response is relevant, complete, and correct for the task.
- Stability and operational tradeoffs: whether results hold across repeated runs and whether the configuration fits the application’s operating constraints.
There is no universal performance threshold or single metric that proves a RAG system is good. Set acceptance criteria for the task, use the same evaluation cases when comparing changes, and investigate failures that matter even when aggregate scores look favorable.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsQuick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

