Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To build a support agent that shows where its answers come from, retrieve relevant knowledge-base chunks first, pass those chunks to a model as context, and return the answer with citations tied to the source documents. Cloudflare AI Search provides a managed retrieval route; Cloudflare also documents a more hands-on stack using Workers AI, Vectorize, and D1.

How the answer-and-citation flow works

  1. Ask: A user sends a support question to a Worker endpoint.
  2. Retrieve: The Worker queries the knowledge base and receives matching document chunks.
  3. Generate: Those chunks are supplied to a model as context for its response.
  4. Return: The application sends back the answer alongside citations identifying the retrieved sources.

Cloudflare’s guide explains that “AI Search returns the source chunks it uses to generate an answer.” The citation should identify the document behind a chunk, not merely assert that the answer is correct. Make sources inspectable so users can check the supporting material; source display can also help you diagnose retrieval quality. Cloudflare’s citation guide covers both standard and streaming responses.

Choose a retrieval architecture

Managed route: Cloudflare AI Search

AI Search is the shorter path when you want Cloudflare to handle ingestion, indexing, and querying for support material. The service can index connected websites, R2 buckets, and uploaded documents. Its overview describes automated indexing, custom metadata filtering, hybrid semantic-and-keyword retrieval enabled by default, OCR for scanned PDFs and images, and a built-in MCP endpoint. Query it through a Worker binding and use the returned content as model context. Read Cloudflare’s AI Search overview.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Assemble more of the RAG stack yourself

If you want to work directly with retrieval components, Cloudflare’s tutorial builds a Worker application with Workers AI, Vectorize, and D1. It starts with a project created using npm create cloudflare@latest, then uses Wrangler for development and deployment and adds an AI binding for model access. This route gives you a more self-managed implementation path; it also means assembling more of the ingestion and retrieval system than with managed AI Search. Follow Cloudflare’s RAG tutorial.

Set up the Worker and connect your knowledge base

  1. Create a Worker project. The RAG tutorial uses npm create cloudflare@latest to get started. Follow the tutorial’s project setup for its Workers AI, Vectorize, and D1 path, or configure the AI Search binding for the managed path.
  2. Configure access to AI Search. Cloudflare documents namespace bindings, which can access and manage instances at runtime, as well as instance bindings for a specific instance. Choose based on whether the Worker needs access to multiple instances or just one. Consult the Workers bindings reference for the binding configuration and API.
  3. Connect the support material. For AI Search, connect a website or R2 bucket, or upload documents, and allow indexing to run. For the self-managed tutorial, follow its steps for building the retrieval path with Vectorize and D1.
  4. Develop and deploy with Wrangler. The tutorial uses wrangler dev for local development and wrangler deploy to deploy the Worker. Test the question-to-retrieval-to-answer path before exposing the endpoint to users.

Return citations users can actually inspect

Use each retrieved chunk’s item.key as the source identifier. It typically contains a filename or URL. Return a human-readable citation—such as a document title linked to its original URL—along with a relevant snippet or useful metadata. If multiple retrieved chunks have the same key, group them into one document citation instead of showing duplicate source entries. The bindings documentation describes returned chunk information that can include the source key, item timestamp, custom metadata, text, and relevance scoring fields.

Preserve the chunk text or a useful excerpt so the interface can show why the source was retrieved, but make clear that the excerpt is evidence to inspect, not a guarantee that every sentence in the generated answer follows from it. A citation points to the material used for generation; the application still needs to present that material clearly enough for verification.

Tune retrieval only when you have a reason

AI Search’s hybrid semantic-and-keyword retrieval is enabled by default, so begin by checking whether it finds the right support passages for real user questions. If retrieval results are poor, inspect the returned chunks and source metadata before changing settings; the problem may be the indexed material or the query rather than ranking alone.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reranking is disabled by default. Cloudflare says it may improve ordering on large or noisy datasets, but enabling it adds a request step that can increase latency; the documentation does not quantify that increase. Compare results on your own support questions before deciding whether the ordering improvement is worth the additional query step. See the reranking configuration guide.

Choose bindings, tuning, and model providers

Decision Option one Option two What to weigh
Retrieval architecture Managed AI Search Workers AI, Vectorize, and D1 tutorial stack How much ingestion, indexing, and retrieval infrastructure you want Cloudflare to manage versus assemble and control.
Worker binding Namespace binding Instance binding Use the documented namespace approach for access to multiple instances; use an instance binding when the Worker targets a specific instance.
Retrieval tuning Default hybrid search Add reranking when warranted Whether improved ordering on a large or noisy corpus justifies an extra request step that may increase latency.
Model provider Workers AI Other providers via AI Gateway Provider and model requirements, plus the work of tracking model lifecycle changes and testing replacements.

Cloudflare’s model documentation advises monitoring model lifecycle information and testing replacements as models change. Treat model selection as an operational decision, not a permanent setting: a replacement may require migration work and validation. Review Workers AI model lifecycle information.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What citations do—and do not—establish

A citation lets a reader inspect a retrieved document that informed the answer. It does not, by itself, prove the answer is complete, current, or correct. Retrieval can surface an irrelevant passage, and a generated response can overstate what its sources support. Keep the source identifier and supporting text available, and evaluate responses against representative support questions before relying on them in a customer-facing workflow.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.