Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

LLM retrieval works as a pipeline: prepare documents, split them into searchable passages, index the passages, find likely matches, rank the results, and send useful context to the model. An agent can add query planning and tool calls around that pipeline, but it is an optional orchestration choice—not another name for retrieval or RAG.

How does an LLM retrieval pipeline work?

Retrieval-augmented generation (RAG) gives a model relevant material from an external source collection when it answers a query. The collection is prepared ahead of time; the incoming query is processed at answer time. The stages are related, but each solves a different problem.

  1. Prepare sources: clean and format documents so their content can be processed and searched.
  2. Chunk documents: divide longer sources into passages that can be retrieved independently.
  3. Index passages: store searchable text, and—when using vector search—embeddings. Keep useful metadata such as titles, URLs, or filenames with the content.
  4. Retrieve candidates: use the query to find passages that may answer it, through keyword search, vector search, or both.
  5. Rank or combine results: order or fuse candidates so stronger matches are more likely to be selected.
  6. Ground the response: provide the selected passages along with the query in the prompt sent to the model.

This separation matters when answers are poor: the issue might be missing or messy source data, unsuitable chunk boundaries, weak search configuration, or ranking—not necessarily the language model itself.

What does chunking do, and how should you choose chunks?

Chunking turns a document into smaller units that can be matched and returned separately. If the units are too broad for the query, retrieval may bring along irrelevant material; if a passage depends on context left outside its boundaries, the retrieved text may not make sense on its own. Chunk boundaries therefore affect what information can be retrieved together.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no universally correct chunk size or overlap rule established by the sources cited here. The useful choice depends on the documents and the questions readers will ask. For example, a short reference entry and a long narrative section may need different boundaries to preserve meaning. Evaluate chunking against representative queries and inspect whether the returned passages contain enough context to answer them.

What is indexing, and why keep metadata?

An index makes prepared content searchable. In a vector-search system, text is represented by embeddings so a search can find semantically similar passages, including some that do not use the query’s exact wording. An index may also support keyword search and filters, depending on the system.

Metadata keeps a passage connected to its origin. Fields such as a document title, URL, or filename can help identify where retrieved information came from and support more useful citations in generated answers. If provenance matters, preserve these identifiers during preparation and make them available alongside the text the model receives.

How do retrieval, scoring, and reranking differ?

Retrieval finds candidate passages; scoring and ranking order or refine those candidates. The search method determines what counts as a promising match, while configured scoring or reranking can adjust which candidates appear nearer the top. A score is a signal used by a particular system—not a universal probability that a passage is correct or sufficient.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Approach What it is useful for What to watch for
Keyword search Queries where exact terms, identifiers, names, or wording matter. A passage that expresses the same idea with different words may not match as well.
Vector search Finding passages with semantically similar meaning, including paraphrases. Semantic similarity does not establish that a passage answers the question accurately.
Hybrid search Combining keyword and vector results to cover both exact-term and semantic matching. The two result sets need a method for combination and ordering.
Semantic ranking or reranking Refining the order of retrieved candidates according to additional relevance signals. Ranking improves ordering; it cannot make an incomplete source answer a question.

These methods are configurable capabilities, not one fixed recipe. For instance, Azure AI Search documentation describes chunking and vectorization, hybrid queries, semantic ranking, and scoring profiles; Progress documentation discusses rank fusion and reranking. Their details are system-specific, so do not assume a score or ranking formula behaves the same across products.

Where do agents fit?

An agent can plan a query or sequence of steps and use retrieval as one capability among others, including access to multiple sources. This is often called agentic retrieval. It is distinct from classic RAG, in which a more direct retrieval pipeline supplies context for a response.

Microsoft’s Azure AI Search guidance positions agentic retrieval for complex or conversational queries and structured responses, while classic RAG can suit simpler needs or situations where speed, generally available capabilities, or fine-grained pipeline control are priorities. Those are product-specific recommendations, not a rule that every complex query needs an agent. An agent adds orchestration; it does not remove the need for suitable sources, retrieval, ranking, and grounded context.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should you choose an approach?

Start with the query workload and the requirements for the answer rather than choosing a fashionable component first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Exactness: Does the task depend on exact names, codes, or phrases, semantic paraphrases, or both? Consider keyword, vector, or hybrid retrieval accordingly.
  • Query complexity: Is one direct search usually enough, or do questions require planning, conversation context, or multiple sources?
  • Provenance: Must answers point readers back to source documents? Retain identifiers such as titles and URLs in the index.
  • Relevance: Do retrieved candidates need another ranking step, and can you assess that step against real queries?
  • Safety and operations: Account for filters, system instructions, latency, and implementation cost in the actual architecture. Google Cloud’s published RAG architecture, for example, includes safety filters and system instructions; these are architectural choices, not automatic effects of retrieval.

Provider documentation from Microsoft, AWS, and Google describes different implementation paths, but the sources do not provide a comparable benchmark establishing one provider or retrieval design as universally faster, cheaper, or more accurate. Measure those trade-offs on the system and corpus you intend to use.

What to check when retrieval answers are weak

  • The right material is absent: verify that source preparation and indexing included the document or passage the query needs.
  • A passage lacks necessary context: review its chunk boundaries and whether related information was separated.
  • Relevant wording is not being found: inspect the retrieval mode and search configuration; exact-term and semantic queries can behave differently.
  • A relevant candidate is buried: review ranking, fusion, or reranking settings rather than treating the top result as inherently correct.
  • The answer cannot be traced: check that source titles, URLs, filenames, or other identifiers survived indexing and are passed through with retrieved text.

Microsoft Foundry’s RAG guidance likewise identifies chunking, embedding quality, and search configuration as areas to review when retrieval is poor. These checks locate different failure points in the pipeline; changing the prompt alone will not repair missing or badly retrieved source context.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.