Do not pass every retrieved chunk—or blindly take the highest-scoring few. Treat retrieval results as candidates, then select the evidence that answers this query, preserves enough context to interpret it, covers every part of the question, and fits your latency and token budget. If the evidence is incomplete or contradictory, retrieve again or abstain instead of asking the LLM to guess.
What makes a retrieved chunk worth sending?
A retrieval score is a ranking signal, not a calibrated probability that a passage supports an answer. A chunk can be top-ranked yet miss the decisive fact, lose the name or date that gives it meaning, or conflict with another source. Conversely, a passage ranked lower on its own may supply a fact needed to complete the answer.
Separate two questions:
- Relevance: Does the passage concern the user’s question?
- Sufficiency: Does the selected evidence, taken together, contain what is needed to answer every part definitively?
Google Research defines context as sufficient when it contains all information necessary for a definitive answer; context is insufficient when it is missing information, incomplete, inconclusive, or contradictory. That distinction matters because a prompt full of relevant-looking text can still fail to support an answer. Google Research’s sufficient-context analysis also warns that adding context can make a model less likely to abstain appropriately when the evidence is inadequate.
How should you select evidence for the LLM?
Use a staged policy. Keep the original documents and provenance available throughout, so the system can restore surrounding context, cite the source, and investigate disagreements.
#1 Best Overall
- Clarify the information need. Turn conversational follow-ups such as “what about the other plan?” into a standalone query that includes the missing subject. For compound questions, list the distinct facts the answer must establish. NVIDIA’s RAG query-to-answer pipeline documentation describes query rewriting as an optional pipeline step; ACL 2025 research on set selection examines the need to cover multiple facts in multi-hop questions.
- Retrieve a candidate pool before filtering. Semantic search can find paraphrases and conceptual matches; lexical search can catch exact names, identifiers, codes, and phrases. When both matching modes matter, combine and deduplicate their results. Anthropic describes a hybrid of vector search and BM25, while Microsoft recommends hybrid keyword and vector retrieval to improve recall. See Anthropic’s Contextual Retrieval guidance and Microsoft’s Azure AI Search RAG overview.
- Repair missing context. A chunk split from its document may omit the entity, timeframe, or subject needed to interpret it. Preserve source metadata and a route back to neighboring passages. One indexing option is to prepend a concise, document-specific context to each chunk; Anthropic describes this as contextual retrieval. Use it where chunk boundaries actually cause retrieval failures, and evaluate the added indexing and storage work.
- Rerank the candidates against the query. A reranker scores a broader retrieved pool for query-specific fit, after which the system can keep a smaller selection. It is a filtering stage, not proof that the remaining passages are complete or correct. Measure whether reranking improves answer quality or groundedness enough to pay for its added latency and cost; the NVIDIA pipeline documentation presents reranking as one component of a retrieval-to-answer system.
- Check coverage as a set. For each required fact, identify which selected passage supports it. Look for repeated passages that crowd out a complementary fact, missing dates or entity qualifiers, and sources that disagree. Set-wise selection is especially relevant for multi-hop questions: several individually relevant chunks can still fail collectively. The findings in the ACL set-selection paper concern its evaluated multi-hop benchmarks, so test whether the approach helps your own query mix.
- Generate with boundaries—or do not generate yet. Instruct the model to ground claims in the supplied evidence, distinguish supported facts from inference, and surface unresolved conflicts. When coverage is inadequate, retrieve again, ask a clarifying question, or abstain. Do not treat the mere presence of retrieved text as evidence that an answer is justified.
How many chunks should you pass?
There is no universal top-k. The useful count depends on chunk size, corpus structure, question complexity, model context limits, and how much redundancy the retrieval policy produces. Increasing the count can add missing evidence, but it can also bury the useful passages in noise and raise token use, latency, and cost.
Anthropic reports that 20 chunks performed better than 5 or 10 in the configurations it tested, while cautioning that additional context can distract and recommending experiments on the actual use case. That result is not a general setting to copy: the right comparison is between complete retrieval-and-answer policies on representative questions from your corpus. Anthropic’s article also states, “Always run evals.”
Rank #2
Set a starting candidate-pool size and final context budget, then test nearby values rather than assuming that more is better. Measure answer correctness and evidence coverage alongside prompt tokens and end-to-end latency. If a larger pool improves recall but overwhelms generation, use reranking or set-wise selection to narrow it; if narrowing drops necessary facts, broaden retrieval or improve the query and chunk context.
When are hybrid retrieval and reranking worth the extra stages?
Hybrid retrieval is useful when the corpus contains both concepts expressed in varied language and exact terms that matter. A product code, legal clause number, or person’s exact name may favor lexical matching; a paraphrased description may favor semantic search. Microsoft’s Azure guidance recommends hybrid queries, and Anthropic describes combining and deduplicating BM25 and vector results. Neither method removes the need to evaluate whether the candidates cover the question.
Rank #3
Reranking is worth considering when a broader candidate pool contains useful evidence but its initial ordering is not good enough for a constrained generation budget. It can improve which passages are prioritized, but it cannot recover a document that retrieval never returned, restore context that was lost during chunking, or guarantee set-level completeness by itself. Track the added runtime and cost against measured gains on your workload.
Anthropic’s 2024 article reports lower top-20 retrieval failure rates in its evaluated configurations: contextual embeddings reduced the rate from 5.7% to 3.7% (35% lower); contextual embeddings combined with contextual BM25 reduced it from 5.7% to 2.9% (49% lower); and adding reranking to that combination reduced it from 5.7% to 1.9% (67% lower). These are vendor-reported experimental results under Anthropic’s methodology, not forecasts for another corpus. Use them as evidence that these techniques can help, not as a promised production improvement. Anthropic’s methodology and results.
Rank #4
How can you tell whether the selected context is enough?
Evaluate the evidence set, not just retrieval rank or whether the final answer sounds plausible. Build a test set of representative questions, including follow-ups, exact-name or identifier queries, compound questions, and cases where the corpus lacks an answer or contains conflicting information.
- Coverage: Can a reviewer point to evidence for every material part of the expected answer?
- Correctness and groundedness: Are the generated claims supported by the selected passages, with no unsupported additions?
- Retrieval balance: Does the policy find exact matches and semantic paraphrases without letting duplicate evidence displace complementary facts?
- Context integrity: Are source, entity, date, and relevant neighboring explanation retained?
- Efficiency: What are the prompt-token use, latency, and cost for each policy?
- Failure behavior: Does the system notice missing or contradictory evidence and retrieve again, clarify, or abstain?
Google Research reports at least 93% classification accuracy for its optimized prompted sufficient-context autorater on the evaluation described in its May 14, 2025 article. That is a result for the study’s evaluator and test, not a general production guarantee. An automated sufficiency check can help triage cases, but validate it against human judgments on your own questions and sources. Google Research explains its evaluation.
Keep disagreement visible. Two sources may differ because one is newer, applies to another edition or region, or is simply inconsistent. The system should use documented authority and scope rules where available; if the evidence does not resolve the conflict, it should say so rather than quietly selecting one passage.
When should a simple retrieval pipeline become more complex?
A straightforward pipeline—retrieve, optionally rerank, then generate—offers simplicity, speed, and fine-grained control. More elaborate query planning or agentic retrieval may be justified for conversational requests, questions that need multiple searches, or responses that require structured citations. Microsoft distinguishes these operational fits in its Azure AI Search RAG overview. Complexity is not an accuracy feature by itself: adopt it when evaluation shows that the workload needs it and its additional runtime and operational burden are worthwhile.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

