Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A production RAG system over PDFs is a chain of evidence transformations: source file, then extracted text and layout, then chunk, then index entry, then retrieved passage, then generated claim, then citation. A fluent answer is not enough. The system has to keep an auditable path from each claim back to a specific document version, page, and passage, and it has to measure where that path breaks. This article walks through the pipeline stage by stage, with the evidence each stage must keep and the failures that appear when it does not.
Where evidence can be lost
Each stage either preserves the link to the source or quietly drops it. The table below lists what each stage produces, what it must retain, and what goes wrong when that information is lost.
| Stage | What it produces | Evidence to keep | What goes wrong if it is lost |
|---|---|---|---|
| Source ingestion | Document record | Original PDF, stable document ID, version, upload details, approval status | Withdrawn or superseded files stay searchable, and no answer can say which version it used |
| Extraction | Text, reading order, tables, OCR output | Parser and OCR versions, page-level output | Text is lost or misread, and generation cannot restore it |
| Chunking and metadata | Retrieval units | Section hierarchy, page range, document ID and version, chunking version | Citations point to the wrong place or to nothing |
| Embedding and indexing | Vectors and index | Embedding and index build versions | A mix of configurations gives inconsistent results that are hard to reproduce |
| Retrieval | Ranked candidate passages | Query, filters applied, passage IDs | You cannot tell whether the evidence was missing or was present and ignored |
| Generation | Answer with claims | Prompt version, model version, claim-to-passage mapping | A plausible answer with unsupported claims |
| Output validation | Released answer or abstention | Validation result and reason for any abstention | Unsupported text reaches the user or a downstream tool |
How do I build a RAG pipeline over PDFs?
Build the pipeline so that each stage writes its output together with the identifiers of its inputs. The five stages below cover ingestion through retrieval. Citation and evaluation follow in their own sections.
Recommended Free Tools
1. Establish source identity and lifecycle
Keep the original PDF as the authoritative artifact. Never treat extracted text as a replacement for it. Give each document a stable ID that survives re-uploads of the same file, and store the following for every version:
#1 Best Overall
- the source location and the upload path
- who uploaded the document, when, and from what source
- any approval or access metadata the application needs
- the document’s version or last-updated time, and the ingestion time
OWASP’s RAG Security Cheat Sheet recommends recording this kind of provenance, including who uploaded a document, when, from what source, and with what approval. OWASP RAG Security Cheat Sheet
Decide deletion and revocation behavior before launch. When a document is withdrawn or superseded in the source system, its chunks must be removed from the index or filtered out of retrieval on a defined schedule. Otherwise stale or revoked content stays usable after everyone has agreed it should not be.
2. Extract text and structure
Classify each PDF before choosing an extraction path. The GOV.UK AI Insights guidance on RAG systems makes the point directly: “For instance, audio data needs a transcription pipeline to convert the audio data into text, while ingestion of PDF documents or image files requires corresponding preprocessing techniques.” GOV.UK AI Insights: RAG systems
The main classes of input need different handling:
- Embedded-text PDFs can often be read directly, but reading order still needs checking.
- Scanned pages need OCR. Every OCR error becomes an error in every answer that relies on that text.
- Tables and charts lose their meaning when flattened into a stream of words. Text embedded in images needs its own validation.
- Multi-column and footnote-heavy layouts can interleave columns or detach footnotes from the sentences they qualify.
NVIDIA’s RAG documentation shows several configurable extraction options. These are implementation examples, not a recommendation for a particular parser. NVIDIA accuracy and performance settings NVIDIA publishes these pages under a rolling latest path, so confirm any setting against the release you deploy.
Validate extraction before chunking:
- Assemble a hand-picked sample that mirrors the real corpus: scans, multi-column pages, tables, footnotes, diagrams, and any non-Latin text.
- Compare the extracted output with the rendered page, page by page.
- Log each defect class you find, such as dropped text, merged columns, broken table cells, or OCR substitutions.
- Fix or flag the defects before chunking. Defects at this stage cannot be repaired downstream.
The expected result is a defect log for each parser configuration you test. A configuration that looks fine on prose and fails on tables should not be promoted to production.
Rank #2
3. Normalize, chunk, and attach metadata
Keep the section hierarchy with every chunk. A passage about an eligibility rule should carry the heading path it sits under, so that the rule is not separated from the condition that qualifies it. Attach document identity, version, page number, and section location to each retrieval unit. NVIDIA’s documentation describes page number as processing metadata that can be used in retrieval filters and in citations. NVIDIA custom metadata documentation
An illustrative record might look like this. The field names are examples chosen for this article, not a fixed NVIDIA schema:
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall{
"chunk_id": "policy-7f3a-v4-c0182",
"document_id": "policy-7f3a",
"document_version": "4",
"page_start": 12,
"page_end": 13,
"section_path": ["3 Benefits", "3.2 Eligibility"],
"parser_version": "2026-09-a",
"chunking_version": "hier-v2",
"ingested_at": "2026-09-30T08:14:00Z"
}
Chunk size is a trade-off, not a constant. Tune it against a task-specific evaluation set rather than a default:
| Unit size | Retrieval effect | Risk | Common mitigation |
|---|---|---|---|
| Smaller | Sharper match to a specific question | Loses surrounding context, such as a rule separated from its exception | Attach the section path and neighboring chunk IDs |
| Larger | Preserves context within a passage | Dilutes topical focus and increases the size of the retrieved context | Measure answer quality against retrieved context size on your own questions |
A 2026 arXiv preprint offers one quantitative comparison. Its benchmark covered 36 Portuguese administrative documents (1,706 pages, roughly 492,000 words), used a manually curated set of 50 questions, and tested 19 pipeline configurations. The authors report that metadata enrichment and hierarchy-aware chunking contributed more to question-answering accuracy than the choice of conversion framework. That finding is specific to that corpus, language, and evaluation design. It does not establish a universal parser or chunk size. arXiv preprint 2604.04948 For chunk setting guidance, see NVIDIA accuracy and performance settings, which describes the defaults of NVIDIA’s own blueprint rather than a general optimum.
4. Embed and index with versioned configuration
Treat the parser, normalization rules, chunking method, embedding model, and index configuration as one versioned pipeline. Store that version with every chunk. A change to any of them is a pipeline release, not a parameter tweak.
Rank #3
GOV.UK’s guidance states the operational consequence: “During this step, the selection of the underlying embedding method is also crucial, as altering the chunking as well as the embedding strategy necessitates re-indexing all chunks.” GOV.UK AI Insights: RAG systems Plan the rebuild the way you would plan any data migration. Build the new index alongside the old one, evaluate it on the same question set, switch traffic only when the results hold, and keep the previous index available for rollback.
5. Retrieve and assemble evidence
At query time, retrieve candidate passages, apply the user’s authorization constraints, and then rank or filter the results before any context is assembled. Access control belongs inside retrieval. A filter applied after the model has already seen a restricted passage does not protect that content.
Reranking, hybrid search, and query decomposition are available features in NVIDIA’s RAG documentation. They are not compulsory stages. Enable each one only when your evaluation shows it helps the questions users actually ask. NVIDIA RAG documentation index
Give every passage a stable identifier that links to its document version and location. The answer can then cite what was actually placed in the context window, rather than what the system later claims it used.
How do I make a RAG system cite its sources?
A citation is only as reliable as the mapping between a claim and a passage. A citation that names a relevant document is weaker than one that names the passage supporting the exact claim, because the first kind lets a reader find the document but not the sentence. Build the citation path in four steps:
- Pass each retrieved passage to the model with its passage ID, document ID, document version, page, and section path.
- Instruct the model to ground factual claims in the supplied passages and to tag each claim with the IDs of the passages that support it.
- After generation, validate each tag. The ID must exist in the retrieved set, the passage text must support the claim, and the document version must still be current.
- If no retrieved passage supports a claim, remove that claim. If no claims remain, return an insufficient-evidence response instead of an answer.
The last step is the abstention behavior. When retrieval does not support an answer, the system should say so. A fluent answer with no supporting passage is the failure this design exists to prevent. OWASP’s guidance treats generation and output validation as separate control points for this reason. OWASP RAG Security Cheat Sheet
How do I know which stage failed?
A single accuracy score cannot tell you whether the retriever missed the evidence or the model ignored it. Score the stages separately. NVIDIA’s evaluation documentation lists answer accuracy, context relevancy, response groundedness, and context recall as distinct measures. NVIDIA RAG evaluation documentation
| Measure | Stage it diagnoses | What a low score suggests |
|---|---|---|
| Context recall | Retrieval | The evidence needed for the answer never reached the context |
| Context relevancy | Retrieval and ranking | Retrieved passages are off-topic or padded with material that does not bear on the question |
| Response groundedness | Generation | Claims in the answer are not supported by the retrieved context |
| Answer accuracy | End to end | The final answer disagrees with the expected answer |
Build the evaluation set from real user questions, with expected answers or expected evidence wherever you can supply them. Add checks that the four measures above do not cover: citation correctness, abstention behavior, stale-document exposure, permission enforcement, latency, and cost. Which of these you need depends on the application.
Re-run the evaluation after any change to parsing, OCR, chunking, embeddings, retrieval, reranking, prompts, or model versions. Keep a trace for every answer that records the query, the filters applied, the document versions and passage IDs retrieved, and the prompt and model versions used. Without that trace, a regression cannot be explained.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteTreat documents as a security boundary
PDFs and extracted text are untrusted input. OWASP’s guidance covers risks and controls across ingestion, embedding generation, vector storage, retrieval, response generation, output validation, and downstream agent integration. OWASP RAG Security Cheat Sheet In practice, that means:
Best Value
- Attach access metadata to each document and enforce it in retrieval, before content reaches the model. Enforce tenant isolation in storage as well as in query filters.
- Treat retrieved text as data, never as instructions. A PDF can contain text that does not appear in ordinary rendering but is picked up by extraction, and that text can try to steer the model.
- Validate generated output before it reaches a user or a downstream tool. This matters most when an agent acts on the answer.
- Keep the provenance record from step 1 for every upload, so that a malicious or erroneous document can be traced and removed.
Choosing components without a benchmark
Published guidance does not provide a neutral, current vendor-by-vendor comparison that would justify naming one parser, embedding model, vector store, or chunk size. Choose components by measuring them on your own corpus against the axes below.
| Axis | What to measure on your corpus |
|---|---|
| Extraction fidelity | Embedded text, scans, tables, charts, multi-column layout, and the languages you actually serve |
| Structure and citation metadata | Whether headings, page locations, and version identifiers survive into retrieval units |
| Retrieval and groundedness | Context recall, context relevancy, and response groundedness on representative questions |
| Lifecycle operations | Update, deletion, re-index, and rollback behavior |
| Security and auditability | Access control, tenant isolation, trace completeness, and resistance to untrusted document content |
| Operating cost | Latency, cost, and operational effort under your expected workload |
Component categories you will evaluate include PDF parsing and OCR tools, managed RAG services, vector retrieval, and reference implementations such as NVIDIA’s RAG Blueprint. A reference implementation shows how the stages connect and which settings exist. It does not replace testing against your own documents.
When an answer is wrong, where to look
Use the symptom to choose the first stage to inspect, then check the stage’s recorded output rather than the final answer alone.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
| Symptom | Likely stage | First check |
|---|---|---|
| The right document never appears in results | Extraction or chunking | Find the answer text in the extracted output for that page, then check whether any chunk contains it |
| The passage is retrieved but the answer ignores it | Generation | Confirm the passage was inside the context window and review the groundedness score |
| The citation names the right document but the wrong page | Metadata | Trace the page number from the parser output through to the chunk record |
| The answer relies on a superseded policy | Lifecycle | Check the document’s version status and the index build it came from |
| A figure from a table is wrong while the citation is correct | Extraction | Compare the table cell in the rendered page with the extracted cell |
| The answer is confident but no passage supports it | Output validation | Confirm the claim-to-passage check runs and that abstention triggers when it fails |
A pipeline that keeps these records can be debugged the same way every time. That consistency is the practical meaning of production-ready.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

