Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A minimal retrieval-augmented generation (RAG) system has four jobs: split source documents into traceable chunks, embed those chunks, retrieve relevant passages for a question, and generate an answer that cites the passages it used. The example below keeps these steps visible in Python; you can replace its in-memory vector search with a hosted vector store when you need persistence or managed retrieval.
What a minimal RAG pipeline does
RAG adds retrieved source material to a language model’s input at answer time. The model does not search your files by itself: your application prepares the material, finds relevant passages for each question, and supplies them to the model.
- Parse: extract text and useful location details from each source.
- Chunk: divide the text into passages small enough to retrieve usefully.
- Embed: turn each passage into a numeric vector representing its semantic content.
- Retrieve: embed the question and rank stored passages by similarity.
- Generate and cite: ask a language model to answer from the selected passages, then map citations back to their source locations.
The vector is only one part of a record. Keep each chunk’s text and provenance—such as document ID, title, URL or filename, section, page, and offsets—together so a match can become a useful citation.
How to represent documents and chunks
Use stable identifiers and preserve enough information to find the original passage. For example, a parsed document can look like this:
#1 Best Overall
document = {
"id": "handbook-01",
"title": "Employee Handbook",
"source": "handbook.pdf",
"text": extracted_text,
}
Each chunk should carry its own identifier and source location. The exact fields depend on your parser, but a useful shape is:
chunk = {
"id": "handbook-01:section-3:chunk-2",
"document_id": "handbook-01",
"title": "Employee Handbook",
"source": "handbook.pdf",
"section": "Leave policy",
"start_char": 4200,
"end_char": 4870,
"text": "...",
}
Parse each file format deliberately and surface parse failures rather than silently indexing empty or malformed text. Normalize whitespace, but do not discard headings, table labels, page numbers, or other structure needed to understand the passage. Capture locations before chunking when possible; reconstructing them later is error-prone.
How to chunk documents for RAG
Start with meaningful boundaries such as headings and paragraphs, then apply a maximum size so a very long section does not become one retrieval unit. A chunk that is too large can dilute a focused match; one that is too small can omit the context needed to interpret a sentence. Overlap can help preserve continuity across boundaries, but it duplicates text in storage and can cause repeated material to appear in retrieved context.
Rank #2
There is no universally correct chunk size established by the sources here. OpenAI’s managed vector-store file API documents an automatic strategy with a maximum chunk size of 800 tokens and 400-token overlap. Its static strategy permits 100–4,096 maximum chunk tokens, with overlap no greater than half the maximum chunk size. These are OpenAI API settings and constraints, not general RAG recommendations. See the Vector store files API reference.
For a small local implementation, a simple replaceable chunker can split on blank lines and combine paragraphs until a character ceiling is reached. Character counts are only a rough proxy for tokens, so production code should use a tokenizer appropriate to its embedding model.
def chunk_paragraphs(text, max_chars=1200):
paragraphs = [p.strip() for p in text.split("nn") if p.strip()]
chunks = []
current = []
length = 0
for paragraph in paragraphs:
if current and length + len(paragraph) + 2 > max_chars:
chunks.append("nn".join(current))
current = []
length = 0
current.append(paragraph)
length += len(paragraph) + 2
if current:
chunks.append("nn".join(current))
return chunks
This basic function does not preserve offsets or metadata by itself. In a real indexer, assign each resulting chunk an ID and location as you create it. Test chunk boundaries against representative questions with known answers; adjust the strategy if the relevant fact is routinely split from its explanation or surrounded by unrelated text.
How to embed and store chunks
An embedding API accepts text and returns a vector. The OpenAI embeddings guide demonstrates Python with client.embeddings.create(input=..., model="text-embedding-3-small") and explains that vectors can be saved in a vector database for later retrieval. Its current documentation lists default dimensions of 1,536 for text-embedding-3-small and 3,072 for text-embedding-3-large, with an 8,192-token maximum input for both listed models. These are provider specifications and can change; consult the OpenAI embeddings guide for current details.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
response = client.embeddings.create(
input=[chunk["text"] for chunk in chunks],
model="text-embedding-3-small",
)
for chunk, item in zip(chunks, response.data):
chunk["embedding"] = item.embedding
Use the same embedding model for indexed chunks and incoming questions, and keep the corresponding vector dimensions consistent. Store the chunk text or a dependable reference to it alongside the vector and provenance metadata. A local in-memory list is convenient for learning and small examples, but it disappears when the process ends; persistence, metadata filtering, updates, and larger indexes require additional storage and operational choices.
How to retrieve relevant context
For each question, create a query embedding and compare it with stored chunk vectors. Cosine similarity is a common ranking method; OpenAI’s embeddings guide recommends it and notes that its embeddings are unit-normalized. For a small local collection, a direct implementation makes the mechanics visible:
import numpy as np
def cosine_similarity(a, b):
a = np.asarray(a)
b = np.asarray(b)
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
def retrieve(query_vector, chunks, limit=5):
ranked = sorted(
chunks,
key=lambda chunk: cosine_similarity(query_vector, chunk["embedding"]),
reverse=True,
)
return ranked[:limit]
For real workloads, use a vector-store search operation rather than repeatedly comparing every vector in application code. OpenAI’s retrieval guide demonstrates searching a vector store with a natural-language query.
Retrieve a candidate set, inspect its relevance, then select the passages that fit the model’s context. Similarity scores rank candidates; they do not prove that a passage answers the question or that the answer is correct. Keyword or hybrid retrieval can be useful for exact names, IDs, dates, and rare terms, but its configuration should be evaluated for your corpus rather than assumed to improve every search.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →How to generate a grounded answer
Pass the question and selected chunks to a generation model in structured form. Instruct it to rely on the supplied evidence, say when the evidence is insufficient, and attach source identifiers to factual claims. Keep the chunk IDs and citation metadata available outside the prompt; do not flatten everything into untraceable text.
Best Value
context = [
{
"id": chunk["id"],
"source": chunk["source"],
"section": chunk["section"],
"text": chunk["text"],
}
for chunk in retrieved
]
# Send the question and context to your chosen generation model.
# Require citations to use only IDs present in context.
The model’s answer should not be trusted to invent valid source references. Treat citation IDs as structured output to validate, not as free-form links written by the model.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to render and validate citations
Resolve every cited chunk ID against the retrieved records, then use that record’s metadata to show a source link or filename and a location such as section or page. Put citations beside the claims they support. If an ID is missing, unknown, or not among the retrieved chunks, reject it or omit the unsupported claim; when no passage supports an answer, return an explicit insufficient-evidence response.
OpenAI’s hosted file-search documentation describes responses that include file citations. A custom pipeline still needs to map its own chunk IDs to original documents and render citations in its own answer format. See OpenAI file search.
When to use local code or managed retrieval
A local implementation exposes parsing, chunking, vector math, and citation mapping, which makes the pipeline easier to inspect and customize. It also leaves you responsible for storage, indexing, updates, and operational reliability. A managed vector store can automate more infrastructure and retrieval work, but introduces provider-specific interfaces and data-handling considerations and may expose fewer implementation details.
These are trade-offs, not evidence that one approach is universally better. Compare both approaches on the same representative questions, checking retrieval relevance and whether every citation resolves to the correct passage. Also test realistic corpus and query volumes and review current pricing before deciding; no comparative quality, latency, or cost benchmark is established here.
Quick Recap
What to test before relying on the system
- Confirm each source was parsed successfully and that its text and location metadata are intact.
- Use questions with known supporting passages to check chunk boundaries and retrieval ranking.
- Verify the query embedding uses the same model as the indexed chunk embeddings.
- Check that retrieved context contains enough evidence without flooding the prompt with irrelevant or duplicate passages.
- Test unsupported questions and ensure the answer does not fabricate facts or citations.
- Resolve every displayed citation to the intended source and passage.
- For a hosted API, recheck current model limits, chunking behavior, and response fields in its documentation.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

