Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To build a support agent that shows where its answers come from, retrieve relevant knowledge-base chunks first, pass those chunks to a model as context, and return the answer with citations tied to the source documents. Cloudflare AI Search provides a managed retrieval route; Cloudflare also documents a more hands-on stack using Workers AI, Vectorize, and D1.
How the answer-and-citation flow works
- Ask: A user sends a support question to a Worker endpoint.
- Retrieve: The Worker queries the knowledge base and receives matching document chunks.
- Generate: Those chunks are supplied to a model as context for its response.
- Return: The application sends back the answer alongside citations identifying the retrieved sources.
Cloudflare’s guide explains that “AI Search returns the source chunks it uses to generate an answer.” The citation should identify the document behind a chunk, not merely assert that the answer is correct. Make sources inspectable so users can check the supporting material; source display can also help you diagnose retrieval quality. Cloudflare’s citation guide covers both standard and streaming responses.
Choose a retrieval architecture
Managed route: Cloudflare AI Search
AI Search is the shorter path when you want Cloudflare to handle ingestion, indexing, and querying for support material. The service can index connected websites, R2 buckets, and uploaded documents. Its overview describes automated indexing, custom metadata filtering, hybrid semantic-and-keyword retrieval enabled by default, OCR for scanned PDFs and images, and a built-in MCP endpoint. Query it through a Worker binding and use the returned content as model context. Read Cloudflare’s AI Search overview.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Assemble more of the RAG stack yourself
If you want to work directly with retrieval components, Cloudflare’s tutorial builds a Worker application with Workers AI, Vectorize, and D1. It starts with a project created using npm create cloudflare@latest, then uses Wrangler for development and deployment and adds an AI binding for model access. This route gives you a more self-managed implementation path; it also means assembling more of the ingestion and retrieval system than with managed AI Search. Follow Cloudflare’s RAG tutorial.
#1 Best Overall
Set up the Worker and connect your knowledge base
- Create a Worker project. The RAG tutorial uses
npm create cloudflare@latestto get started. Follow the tutorial’s project setup for its Workers AI, Vectorize, and D1 path, or configure the AI Search binding for the managed path. - Configure access to AI Search. Cloudflare documents namespace bindings, which can access and manage instances at runtime, as well as instance bindings for a specific instance. Choose based on whether the Worker needs access to multiple instances or just one. Consult the Workers bindings reference for the binding configuration and API.
- Connect the support material. For AI Search, connect a website or R2 bucket, or upload documents, and allow indexing to run. For the self-managed tutorial, follow its steps for building the retrieval path with Vectorize and D1.
- Develop and deploy with Wrangler. The tutorial uses
wrangler devfor local development andwrangler deployto deploy the Worker. Test the question-to-retrieval-to-answer path before exposing the endpoint to users.
Return citations users can actually inspect
Use each retrieved chunk’s item.key as the source identifier. It typically contains a filename or URL. Return a human-readable citation—such as a document title linked to its original URL—along with a relevant snippet or useful metadata. If multiple retrieved chunks have the same key, group them into one document citation instead of showing duplicate source entries. The bindings documentation describes returned chunk information that can include the source key, item timestamp, custom metadata, text, and relevance scoring fields.
Preserve the chunk text or a useful excerpt so the interface can show why the source was retrieved, but make clear that the excerpt is evidence to inspect, not a guarantee that every sentence in the generated answer follows from it. A citation points to the material used for generation; the application still needs to present that material clearly enough for verification.
Rank #2
Tune retrieval only when you have a reason
AI Search’s hybrid semantic-and-keyword retrieval is enabled by default, so begin by checking whether it finds the right support passages for real user questions. If retrieval results are poor, inspect the returned chunks and source metadata before changing settings; the problem may be the indexed material or the query rather than ranking alone.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchReranking is disabled by default. Cloudflare says it may improve ordering on large or noisy datasets, but enabling it adds a request step that can increase latency; the documentation does not quantify that increase. Compare results on your own support questions before deciding whether the ordering improvement is worth the additional query step. See the reranking configuration guide.
Rank #3
Choose bindings, tuning, and model providers
| Decision | Option one | Option two | What to weigh |
|---|---|---|---|
| Retrieval architecture | Managed AI Search | Workers AI, Vectorize, and D1 tutorial stack | How much ingestion, indexing, and retrieval infrastructure you want Cloudflare to manage versus assemble and control. |
| Worker binding | Namespace binding | Instance binding | Use the documented namespace approach for access to multiple instances; use an instance binding when the Worker targets a specific instance. |
| Retrieval tuning | Default hybrid search | Add reranking when warranted | Whether improved ordering on a large or noisy corpus justifies an extra request step that may increase latency. |
| Model provider | Workers AI | Other providers via AI Gateway | Provider and model requirements, plus the work of tracking model lifecycle changes and testing replacements. |
Cloudflare’s model documentation advises monitoring model lifecycle information and testing replacements as models change. Treat model selection as an operational decision, not a permanent setting: a replacement may require migration work and validation. Review Workers AI model lifecycle information.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What citations do—and do not—establish
A citation lets a reader inspect a retrieved document that informed the answer. It does not, by itself, prove the answer is complete, current, or correct. Retrieval can surface an irrelevant passage, and a generated response can overstate what its sources support. Keep the source identifier and supporting text available, and evaluate responses against representative support questions before relying on them in a customer-facing workflow.
Quick Recap
Best Value
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

