Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cloudflare Workers AI runs the model inference; AI Gateway sits in front of requests to add analytics, logging, caching, rate limits, retries, and fallback controls. For a chat application, you can call Workers AI from a Cloudflare Worker using a binding, or send requests to Cloudflare’s REST API. Choose an endpoint that supports both your desired request schema and the selected model, and treat gateway caching and rate limits as operational controls—not conversation memory or a complete abuse-prevention system.

How the pieces fit together

Cloudflare describes AI Gateway as a visibility and control layer for AI applications, while Workers AI provides model inference on Cloudflare’s serverless GPU infrastructure. Gateway can also work with external providers, including OpenAI, Anthropic, and Google. Cloudflare says AI Gateway is available on all plans and describes its core features as free; Workers AI is available on Free and Paid Workers plans. Its overview lists a catalog of 50+ open-source models, a vendor-reported catalog figure rather than a comparison of model quality or latency.

A typical request flow is: a user sends a message to your application; your application constructs a model request; the request goes through a Worker binding or Cloudflare’s REST API with an AI Gateway ID; Workers AI runs the selected model; and the response returns to the application. Gateway analytics and controls can help operators monitor that flow, but your application remains responsible for validating input and output, handling errors, protecting sensitive data, and deciding how conversation history is managed.

Choose a Worker binding or the REST API

Integration Where the inference call runs When it fits Important setup
Worker binding Inside your Cloudflare Worker application You already run application logic on Workers and want the inference call integrated with it. Call env.AI.run(model, input, options) and include a gateway object identifying an existing gateway. Binding options can include skipCache and cacheTtl.
REST API From an application or service making an HTTP request to a Cloudflare account AI endpoint You need an HTTP integration, or want to route to Workers AI or supported third-party providers through the API. For Workers AI, specify the @cf/author/model identifier and send the cf-aig-gateway-id header. Requests to /accounts/{account_id}/ai/* require a Cloudflare API token with Account > Workers AI > Read permission.

Gateway configuration endpoints have separate AI Gateway permissions; do not assume the Workers AI permission alone authorizes gateway administration. The binding and REST API have different application and authentication workflows, so select based on where your application runs and how you manage credentials.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which endpoint should a chat application use?

Cloudflare documents several endpoint schemas, but they are not interchangeable for every model. The model catalog and its supported endpoint combinations can change, so check the current compatibility information for the model you intend to call.

Endpoint or route Use Compatibility note
POST /ai/v1/chat/completions OpenAI chat-completions-compatible chat requests Confirm that the selected Workers AI model supports the endpoint.
POST /ai/v1/responses Agentic workflows using the Responses API Workers AI support depends on the model.
POST /ai/v1/messages Requests using Anthropic’s Messages schema Does not support Workers AI models.
/ai/run Workers AI model inference using the model’s input schema Use the schema documented for the selected model.

For example, a REST chat request can target POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1/chat/completions, use a Workers AI model identifier such as @cf/meta/llama-3.1-8b-instruct, and include cf-aig-gateway-id: YOUR_GATEWAY_ID. This illustrates the endpoint shape and headers; verify that the model identifier remains available and supports chat completions before deployment. Authenticate the account AI request with a token that has Account > Workers AI > Read permission. The exact request body depends on the endpoint’s schema and the model’s current support.

What AI Gateway adds—and what it does not

AI Gateway can expose request counts, token use, costs, and errors, and offers request controls such as rate limiting, retries, and model fallback. Those controls help operate an AI application; their presence alone does not establish that the application is safe, reliable, or inexpensive. Build application-level validation, privacy review, prompt handling, and error handling around them.

Observability and logging

Use analytics to track how the application is being used and to investigate errors and cost patterns. Logging availability and pricing depend on the account’s gateway creation cohort. Cloudflare’s pricing page says accounts whose first gateway was created on or after September 24, 2026 follow Workers Logs pricing and retention; existing customers use the documented legacy limits. Consult the current pricing page for the applicable account path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retries and fallbacks

Retries and fallback models can help shape behavior when a request fails, but they should be configured with application needs in mind. A retry may add another inference request, and a fallback may have different capabilities or output behavior. Decide what errors are retryable, how many attempts are acceptable, and whether the application can safely use an alternate model.

Rate limiting and user quotas

AI Gateway rate limiting lets an operator set a request count over an interval and choose a fixed or sliding window. Once the configured limit is exceeded, the gateway returns HTTP 429 and does not process the request. Use this as one layer in a broader quota design: enforce user- or tenant-level allowances in the application, and ensure client retries do not amplify traffic during a 429 response.

When caching helps—and when it does not

AI Gateway response caching is disabled by default. Cloudflare currently documents caching for text and image responses and says a result is served only for an identical request. Its default cache key combines provider, endpoint, model, provider authentication header, and the full request body. A changed message, conversation history, or model parameter therefore produces a different cache entry.

That makes gateway response caching a better fit for repeated, stable prompts—such as a constrained support flow with a limited set of choices—than for open-ended chat, where each turn and its history commonly change. It is not conversational memory, and Cloudflare’s documentation describes semantic caching as planned future work rather than a current capability.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Workers AI also documents a separate prompt or prefix caching feature for select models. It can reuse a shared input prefix; Cloudflare advises placing static prompt material first and using session affinity to improve the chance of routing to an instance holding cached tensors. This model-level inference optimization is distinct from AI Gateway’s exact-request response cache.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Limits and pricing to check before launch

Gateway limits and Workers AI inference limits are separate. The following figures are Cloudflare’s published documentation values accessed in September 2026; they can change, and some apply only to particular billing arrangements or model groups.

Area Published figure or rule Qualification
AI Gateway cacheable request size 25 MB Gateway limit listed by Cloudflare on September 24, 2026.
AI Gateway maximum cache TTL One month Gateway limit listed by Cloudflare on September 24, 2026.
Unified Billing request rate 200 requests per 60 seconds per gateway Applies to Cloudflare-managed credentials through Unified Billing; does not apply to bring-your-own-key requests. Cloudflare limit listed September 24, 2026.
Workers AI included usage 10,000 Neurons per day at no charge Cloudflare pricing documentation last updated September 17, 2026.
Workers Paid usage above the included allocation $0.011 per 1,000 Neurons Cloudflare pricing documentation last updated September 17, 2026. Some models require a paid billing method.
Default text generation rate 300 requests per minute Workers AI default listed September 17, 2026, except for models requiring the Workers Paid plan.
Paid models covered by the limits table 20 requests per minute on standard billing; 50 requests per minute with prepaid AI Gateway credits Applies to the paid models described on Cloudflare’s limits page; verify model-specific requirements and current prepaid-credit behavior.

Neurons are Cloudflare’s measure of model compute. The pricing documentation also publishes model-level token price tables, so estimate costs against the selected model and expected workload rather than treating every message as having the same price. Gateway core analytics, caching, and rate limiting are described as free on all plans, but logging follows the account cohort rules above.

Deployment checklist

  • Choose a supported model and verify its current endpoint and input schema.
  • Choose a Worker binding or REST API integration based on where the application runs.
  • Attach the existing AI Gateway ID using the documented binding option or REST header.
  • Grant only the permissions needed for inference and gateway administration.
  • Decide whether exact-request caching is useful for the traffic you expect; do not rely on it as chat memory.
  • Set gateway limits alongside application-level user quotas and define sensible behavior for HTTP 429 responses.
  • Review retry and fallback behavior for cost, latency, and model-capability implications.
  • Check current Workers AI model pricing, inference limits, AI Gateway limits, and logging terms before deployment.

Official Cloudflare documentation

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.