iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
An AI agent can give a plainly wrong answer because it never received the information needed to answer correctly. That makes retrieval worth checking—but it does not make retrieval the cause of every failure. Anthropic reported 49% fewer failed retrievals with its Contextual Retrieval method, and 67% fewer when reranking was added. Those are results for Anthropic’s method and evaluation, not a forecast of how much any agent’s task failures will fall.
The practical lesson is to inspect the full context assembled for a failed run before changing models or prompts. If the right policy, identifier, exception, or tool result was missing, investigate how information was found and ranked. If it was present but the agent ignored it, or the failure involved planning or tool use, the fix lies elsewhere.
Why an obvious mistake may start before the model answers
Suppose an agent misses a refund window, selects the wrong SKU, forgets a result returned by a tool earlier in the conversation, or overlooks a customer-specific exception. These are examples of failures that can look like poor reasoning. One possibility is that the agent never had the relevant detail in its usable context.
That can happen when a source was not indexed, a document was split into unhelpful chunks, a search filter excluded the relevant record, or a good match ranked below the limit on how many results were passed to the model. The agent may also receive the right text but fail to use it. These are different problems, even if they produce similarly incorrect answers.
#1 Best Overall
Keep three information paths distinct while diagnosing a run:
- Session state: messages and tool results from the current conversation or task.
- Durable memory: information intended to carry across separate runs, such as a user preference.
- External retrieval: material fetched from a knowledge base or other source for a particular query.
A tool result that vanished from the assembled conversation is not necessarily a knowledge-base search failure. First identify which path was supposed to supply the missing fact.
What the 49% figure does—and does not—say
In its September 19, 2024 engineering article, Anthropic reported 49% fewer failed retrievals with Contextual Retrieval, and 67% fewer when reranking was combined with it. Contextual Retrieval adds chunk-specific explanatory context before creating contextual embeddings and a contextual BM25 index. Anthropic’s figures describe failed retrievals in its evaluation. They do not mean that 49% of agent mistakes are retrieval mistakes, that every agent will improve by that amount, or that overall task accuracy rises by 49%.
Rank #2
The components address how relevant material is represented and found. Contextual embeddings can help a chunk make sense in relation to its source document; BM25 provides lexical matching that can help with exact terms, such as identifiers and technical phrases. Reranking reorders candidates so that the most useful results are more likely to fit within the context limit. Each is an option to test against the system’s own representative tasks, not a guaranteed repair.
How to debug one failed run
Start with a single reproducible failure and preserve the complete trace. The important question is not what the system was meant to retrieve, but what the model actually received.
- Capture the assembled input. Save the exact query, system instructions, previous messages, injected memory, retrieved chunks, tool outputs, and the final context sent to the model.
- Record retrieval details. Keep candidate documents and chunks, scores and ranks, filters, index or version state, and any reranking results. This helps distinguish missing candidates from candidates that were found but not selected.
- Locate the missing fact. Ask: “What did retrieval return?” and “Where did the fact appear in context?” If the fact is absent, investigate acquisition. If it appears in the final context, investigate whether it was clear, current, and usable—and whether the model followed it.
- Classify the failure before changing components. Check whether the problem was retrieval, context assembly, generation, planning, tool invocation, tool execution, or a policy constraint. Do not treat every incorrect answer as a search problem.
- Change one likely cause and rerun the case. Compare the new trace and outcome with the preserved baseline, then check the change on a representative set of tasks so a fix for one example does not hide regressions elsewhere.
Use the failure pattern to choose what to inspect
| What the trace shows | Likely area to investigate |
|---|---|
| The relevant document or chunk is absent from candidates | Ingestion, chunk boundaries, filters, query wording, lexical coverage, or index freshness |
| A relevant candidate exists but falls below the top-k cutoff | Ranking quality, retrieval method, candidate limits, or reranking |
| The right evidence reaches the model but the answer contradicts or ignores it | Context assembly, evidence clarity or position, and whether the generation step follows the evidence |
| A prior tool result or preference is missing | Whether it belongs in session state, durable memory, or external retrieval—and how that state enters the prompt |
| The agent chose the wrong action, called a tool incorrectly, or mishandled its result | Planning, tool invocation, execution, or interpretation of tool output rather than retrieval alone |
| The agent pursued an unsupported request or hit a policy restriction | Intent handling and guardrails; a better retriever may not change the outcome |
Microsoft Research’s March 12, 2026 AgentRx announcement illustrates why this classification matters. Its taxonomy includes plan-adherence failures, invented information, invalid tool calls, misread tool output, intent-plan misalignment, underspecified or unsupported intent, guardrail triggers, and system failures. In a benchmark of 115 manually annotated failed trajectories across τ-bench, Flash, and Magentic-One, Microsoft reported improvements of 23.6% in failure-localization accuracy and 22.9% in root-cause attribution over prompting baselines. Those are Microsoft’s reported benchmark results, not a guarantee for another debugging setup.
When to try lexical search, hybrid retrieval, or reranking
Exact identifiers and technical terms
Semantic similarity alone can be a poor fit for exact strings such as order IDs, SKUs, policy names, and error codes. Ask, “Did exact-match search exist?” Compare semantic search with a hybrid approach that combines semantic retrieval and lexical matching such as BM25. Verify that the exact term is indexed and that filters do not exclude the matching record.
Free tools Windows power users keep installed
One-click scans. No signup required.
Relevant material is found but not surfaced
If the trace contains the needed chunk among candidates but it is ranked too low, investigate ranking and reranking rather than reworking ingestion first. Ask, “Was reranking applied?” Test whether reordering candidates moves the useful evidence into the context passed to the model. A ranking change is only helpful if the candidate set already contains relevant material.
Documents change over time
If an answer relies on obsolete information, check index freshness and version state. A perfect ranker cannot surface an updated source that has not been ingested or indexed, and a stale result can be worse than no result when policies or records have changed.
These interventions have trade-offs in recall, ranking, freshness, latency, complexity, and cost. A useful comparison measures whether the system retrieves the right evidence, whether it fits within the model’s context limit, and whether the final answer uses it—not just whether a search component returns plausible results.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.When retrieval may not be necessary
For a small, stable corpus, compare search against including the material directly in the prompt. Anthropic suggests that a knowledge base under 200,000 tokens—about 500 pages in its example—may be included directly. Treat that as Anthropic’s heuristic, not a universal cutoff. Direct context removes some retrieval plumbing, but the model can still overlook information depending on how context is assembled and used.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For larger or frequently changing collections, retrieval can make the prompt more selective, but it introduces its own failure points: ingestion, chunking, filters, ranking, and freshness. Choose between direct context and retrieval by testing against the corpus size, rate of change, context budget, and task needs.
Best Value
What retrieval benchmarks can tell you
Retrieval quality is measurable, but a retrieval metric is not the same as successful task completion. The July 2026 Agent Retrieval Bench paper by Bowen Qin and Yi Xie evaluates file-level context retrieval for coding agents using 427 samples across 25 repositories, including four positive retrieval task types and selective retrieval. Its authors report that logged trajectories miss every gold file on 27–35% of samples. They also report different leaders by metric: Qwen3-Embedding-4B for weighted MRR, Qwen3-Embedding-8B for weighted Recall@20, and RepoMap for budgeted context yield at 8K tokens. Those results do not establish a universal winner or prove that retrieval alone determines whether a coding agent completes a patch.
For your own agent, evaluate retrieval and end-task outcomes separately. A search configuration can improve candidate recall without fixing planning, tool execution, or the model’s use of evidence. Conversely, a task may succeed despite an imperfect retrieval metric if the missing detail was not needed for that run.
Questions to ask before changing the model
- What did retrieval return, and which results reached the final context?
- Where did the relevant fact appear in that context?
- Did exact-match or lexical search exist for the identifier involved?
- Was reranking applied, and did it change which evidence fit in the context limit?
- Was the missing information supposed to come from this run’s tool output, durable memory, or an external source?
- Did the agent actually see the right thing in a usable form?
- Does the failure still occur when the correct evidence is clearly present?
If the evidence is absent, debug acquisition and ranking. If it is present and the answer remains wrong, move to grounding, planning, tool behavior, or policy. That distinction is more useful than reflexively switching models because an answer looked obvious to a person.

