Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To build a retrieval-augmented generation (RAG) pipeline in Python, collect text from an authorized online source, convert it into documents with traceable metadata, split it into passages, create and store embeddings, then retrieve relevant passages for each question and send them to a language model as context. The model answers from the selected material instead of receiving your entire collection on every request.

How the pipeline fits together

A RAG system has two paths: an ingestion path that prepares and indexes source material, and a query path that retrieves relevant material to support an answer. Ingestion must happen before the system can answer questions about a source.

  1. Ingest: collect authorized text and normalize it into documents.
  2. Prepare: preserve provenance, split documents into retrieval-sized passages, and create embeddings.
  3. Index: store each passage, its vector, and its metadata in a vector store or managed retrieval service.
  4. Answer: retrieve passages relevant to a question and give them to the language model as context.
  5. Refresh: detect source changes and update or remove indexed material as needed.

LlamaIndex describes ingestion in terms of loading, transformation, and indexing, with documents and nodes able to carry metadata. Its Ingestion Pipeline documentation also describes caching and document management. OpenAI’s Retrieval guide describes semantic search over data stored in vector stores.

Build the ingestion path

1. Choose a source you are allowed to collect

Your input could be a website, public document collection, API, or another text source. A loader or connector retrieves the content and turns it into documents. Do not assume that a website permits automated collection: check its terms, relevant robots guidance, copyright or license, authentication requirements, rate limits, and update behavior before building a crawler or connector. Those permissions depend on the specific source.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Normalize text and keep provenance

Normalize encoding and remove navigation, repeated boilerplate, or other noise carefully so that useful context is not discarded. Preserve enough metadata to identify and revisit the original material. A practical document record can include the text, a stable document ID, canonical source URL, title, and retrieval timestamp. These fields are implementation recommendations; LlamaIndex documents the ability to associate metadata with documents and nodes in its Loading Data (Ingestion) guide.

In Python, treat the text and its metadata as one unit throughout the pipeline rather than keeping them in separate lists that can become misaligned. For example, a document record should conceptually pair a passage’s text with its source URL and document ID. That pairing lets the retrieval stage preserve provenance, so an answer can be traced to the material that supported it.

3. Split documents into retrieval passages

Long pages usually need to be divided into smaller passages before indexing. A passage should retain enough surrounding context to make sense on its own and answer likely questions. Use the source structure—such as headings, paragraphs, or sections—when it provides meaningful boundaries; otherwise, consider sentence- or token-based splitting.

Overlap can preserve continuity when a useful idea crosses a passage boundary, but it also stores repeated text and can yield duplicate matches. There is no universally correct chunk size or overlap in the cited documentation. OpenAI’s Retrieval API guide reports a default of 800 tokens per chunk and 400 tokens of overlap for that hosted service, with configurable chunk sizes from 100 to 4,096 tokens. The guide also states that overlap must be non-negative and no greater than half the chunk size. These are service settings, not a benchmark or general tuning prescription; see the Retrieval guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Create embeddings and index the passages

An embedding model converts text into vectors that represent it for similarity search. Create an embedding for each passage, then store the vector alongside the passage and its metadata in a vector index or managed vector store. The embedding is not a replacement for the text: the system needs the associated passage later to provide readable context to the language model.

LlamaIndex’s ingestion documentation describes chaining transformations such as a SentenceSplitter, metadata extraction, and an OpenAIEmbedding, then inserting resulting nodes into a vector store. Its documentation notes that an embedding stage is needed when the pipeline connects to a vector store. Exact APIs can vary by framework release; the retrieved documentation does not identify a release version, so check the documentation for the version you install before copying framework-specific calls. See LlamaIndex’s Ingestion Pipeline guide.

Retrieve passages and generate an answer

At question time, represent the user’s query for search, retrieve the most relevant passages, and send those passages together with the question to the generation model. Do not send the full collection with every request: retrieval selects a smaller set of relevant material for the model to use. Semantic search can return passages that are conceptually similar even when they do not share many exact keywords, as described in OpenAI’s Retrieval documentation. LlamaIndex likewise describes supplying relevant indexed information at query time rather than providing all data on every request in its High-Level Concepts and Question-Answering (RAG) guides.

The context sent to the model should include the retrieved text and, where useful, its source information. Instruct the model to ground its response in that context. Retrieval finds candidate evidence; it does not by itself prove that the evidence is complete, current, or sufficient to answer the question.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose an implementation approach

A framework-managed ingestion pipeline and a hosted retrieval API can both support RAG, but they place different decisions in your hands. The documentation describes capabilities rather than a universal winner or a comparative benchmark.

Decision LlamaIndex ingestion pipeline OpenAI Retrieval API
Connecting source material Loading turns source data into documents; the exact connector and source handling depend on your implementation. Source Managed vector stores are documented; source-specific web crawling or permissions are not established by the Retrieval guide. Source
Parsing, chunking, and metadata Transformations can be chained and customized, and documents or nodes can carry metadata. Source The guide documents configurable chunking for its hosted retrieval service; the chunk defaults and bounds are specific to that service. Source
Embeddings and vector storage The pipeline can include an embedding stage and connect to a vector store. Source OpenAI documents managed vector stores for retrieval. Source
Storage location Remote vector-store integration is documented; a particular store or deployment is implementation-specific. Source Managed vector stores are documented; the guide’s specific storage details depend on the API workflow. Source
Caching and updates Node/transformation caching and document management using document IDs or reference document IDs are documented. A website refresh and deletion policy remains your responsibility. Source A universal source-refresh, deleted-page, or stale-vector policy is not stated in the Retrieval guide. Source
Portability and operational effort Not stated as a comparative measure in the cited LlamaIndex documentation. Source Not stated as a comparative measure in the cited OpenAI documentation. Source
File and token limits Not stated in the cited ingestion-pipeline documentation. Source The Retrieval guide reports a maximum file size of 512 MB and 5,000,000 tokens per file for the API. These are API limits, not recommended document sizes; check the guide for current limits before implementation. Source

Choose based on which parts you need to control. A framework pipeline exposes transformation and vector-store integration choices; a managed retrieval service provides a documented hosted vector-store path and service-specific limits. The cited documentation does not establish that one approach is more accurate, faster, cheaper, or more portable.

Make ingestion repeatable and keep the index fresh

Online material changes, so a production pipeline needs an update strategy in addition to its initial indexing run. Use stable document IDs where the source allows them, record which source page each indexed passage came from, and decide how often to check for changes. When a page changes or disappears, determine how to replace its passages and remove stale vectors. These refresh and deletion policies depend on the source and vector store; there is no single policy established by the cited documentation.

LlamaIndex documents caching of node/transformation combinations and document management that can use document IDs to identify duplicates. Those capabilities can reduce needless reprocessing, but they do not determine when a website has changed or how deleted pages should be handled. See the Ingestion Pipeline guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.