Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI agents cannot assume that a website will admit them just because a person can open it in a browser. Access depends on what the agent intends to do—find pages for search, fetch content for a user’s live task, or collect material for model training—and on whether the site can identify the agent and enforce its own rules.

For site owners, the challenge is to make those distinctions without sacrificing useful referrals, content control, security, privacy, or a viable business model. For users, the practical question is how an agent can prove it is acting legitimately and obtain permission rather than simply imitate a browser.

How do I get AI agents to access websites?

There is no universal “allow AI agents” switch across the web. A site owner decides what to admit, and controls differ by hosting provider and site. An agent needs to use the site’s supported access path, identify itself where required, and respect the site’s technical and policy controls. A user generally cannot override a site’s decision from the agent’s side.

The most useful starting point is to ask which activity the site is deciding about. Cloudflare, for example, separates Search, Agent, and Training behavior. These are Cloudflare’s product categories, not a universal web standard, and one bot may perform more than one kind of activity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ESP32-S3 4.2inch RLCD Development Board, 300 x 400, E-Paper-Like Screen, Supports Wi-Fi & BLE Dual-Mode Communication and AI Voice Interaction, Temperature & Humidity Monitoring, DIY
  • E-Paper-Like Display: 4.2-inch fully reflective RLCD screen (300×400 resolution), low power consumption, no backlight, faster refresh rate, providing an eye-friendly reading experience similar to an e-ink screen.
  • High-Performance Processor: Equipped with an ESP32-S3 dual-core processor (240MHz), supporting 2.4GHz Wi-Fi and Bluetooth 5 (LE) , built-in antenna, easily enabling IoT connectivity and AI applications.
  • Supports AI Voice Interaction: Integrated with an SHTC3 high-precision temperature and humidity sensor and a dual-microphone array (supporting noise reduction/echo cancellation), accurately achieving voice recognition and AI voice interaction, compatible with Xiaozhi AI and large models such as Doubao/DeepSeek/GPT.
  • Long Batt Life and Strong Expandability: Supports 186-50 Li Batt power + R-T-C backup Batt, Micro SD card slot for data storage, and reserved rich interfaces such as UART/I2C/GPIO for easy expansion of DIY projects. (Note: This version doesn't include 186-50 Li Batt)
  • Suitable for DIY Creative Projects and Prototype Development: It can be used to create electronic calendars, smart desktop ornaments, AI intelligent agents, etc., taking into account learning, development and practical application.
Activity What it means Why a site may treat it differently
Search Crawling and indexing content so it can help answer questions later. Search discovery can send people to a publisher, but crawling still consumes resources and may expose content to services the publisher does not want to support.
Agent Real-time automated activity on a person’s behalf, such as a chat agent fetching a page or a browser-use agent interacting with a site. A live task may provide direct utility to a user, but the site must consider authentication, abuse, privacy, and whether the agent may take actions.
Training Collecting content for model training or fine-tuning. A publisher may want discoverability or task-based access while refusing use of its content for training.

Cloudflare’s documentation says that, from September 15, 2026, its defaults for new domains block bots classified as Training or Agent on pages displaying ads while leaving Search allowed. It also documents choices to block AI bots on all pages, block them on ad-displaying pages, or allow them. These are Cloudflare-specific defaults and controls—not rules that automatically apply to every website or CDN. See the Cloudflare documentation for the provider’s current behavior.

What an agent or its user can do

  • Check whether the site offers an API, an agent-compatible interface, or another authorized access method, and use that instead of assuming ordinary browser access extends to automation.
  • Where a site or access provider supports bot verification, use the supported identity mechanism rather than relying on a user-agent label that can be copied.
  • Follow the site’s terms and prompts for login, consent, payment, or human confirmation. Access permission does not automatically authorize every action an agent might attempt.
  • If access is denied, use the site directly or ask the operator about permitted access. A crawler instruction file is not a way for a visitor to grant themselves permission.

Why are websites blocking AI agents?

A website has to balance the visitor’s benefit against the cost and risk of the request. Search indexing may create discovery and referrals; a live agent may help a user complete a task; training collection may provide value to a model developer without an equivalent visit or payment to the publisher. From the site’s perspective, requests that look similar at the network level can therefore have materially different purposes.

Other considerations include bandwidth and compute, repeated fetching of unchanged pages, automated abuse, content licensing and control, and whether a site can distinguish a legitimate agent acting for a user from a bot pretending to be one. These are trade-offs, not proof that all publishers share the same policy or that every agent request is harmful.

Rank #2
GeeekPi EmbodiQ AI Starter Kit for Arduino UNO Q – 4GB RAM, 32GB eMMC, AI Agent HAT, Soil Moisture & Raindrop Sensors, Servo, Acrylic Mount – Natural Language Control
  • Talk to Your Hardware – Control sensors, servos, buzzers, and OLED displays using natural language. No complex coding required – just tell the AI what you want to do
  • Powerful AI Agent Onboard – Built around UNO Q with 4GB RAM and 32GB eMMC storage. Runs the EmbodiQ AI Agent HAT, enabling real-time reasoning and multi-step task execution with conditional logic
  • Versatile Sensor Suite – Includes soil moisture sensor, raindrop sensor, 9g servo motor, and OLED output. Perfect for smart gardening, weather stations, robotics, and automation projects
  • Flexible AI Provider Support – Works with OpenAI, OpenRouter, MiniMax, and any OpenAI-compatible API. Choose your preferred model and switch easily via the web-based interface or terminal REPL
  • Dual‑Architecture & Ready to Use – Python + Arduino co-processing ensures responsive performance. Comes with acrylic mounting bracket for tidy assembly – ideal for makers, educators, and AI enthusiasts

Cloudflare reported in 2026 that over 50% of AI crawler traffic it observed was spent re-fetching unchanged pages. That figure is Cloudflare’s own platform observation, not an industry-wide measurement. It illustrates why freshness and efficient update mechanisms matter to operators, but does not establish how much any particular site spends on crawling.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What does robots.txt do—and what does it not do?

A robots.txt file lets a site publish instructions for crawlers. Common directives include User-agent to name a crawler or group, Disallow and Allow to indicate paths, and Sitemap to point to a sitemap. The IETF formalized the Robots Exclusion Protocol in RFC 9309, as Cloudflare explains in its overview of managed robots.txt and AI-content controls.

The critical limitation is that robots.txt is an instruction, not an access-control boundary. Compliance is voluntary: a crawler can ignore it, and the file does not authenticate a bot or prevent a request. A site that needs to deny access or admit only particular visitors must use enforcement at the application, server, authentication, or security-infrastructure layer as well.

Rank #3
ESP32-C6 2.16inch AMOLED Touch Screen Display Development Board, 480×480
  • High-Performance RISC-V Core and Tri-Mode Wireless Communication---Equipped with an ESP32-C6 32-bit RISC-V processor with a 160MHz clock speed, it features 512KB HP SRAM, 16KB LP SRAM, 320KB ROM, and an external 16MB Flash memory. It supports Wi-Fi 6, Bluetooth 5, and IEEE 802.15.4 (Zigbee 3.0 and Thread), and includes an onboard antenna for excellent RF performance.
  • 2.16-inch AMOLED High-Definition Touchscreen---Features a 2.16-inch capacitive AMOLED touchscreen with a 480×480 resolution and 16.7 million colors. It utilizes a CO5300 driver chip (QSPI interface) and a CST9220 touch chip (I2C interface), minimizing pin usage. AMOLED offers high contrast, wide viewing angles, rich colors, fast response, and a slim, low-power design.
  • AI Voice Dialogue and Sensing Functionality---Designed specifically for the development and functional verification of AI voice dialogue intelligent agent prototypes, it features onboard dual microphones and an audio codec chip, supporting Xiaozhi AI and DeepSeek. The QMI8658 six-axis IMU (3-axis accelerometer, 3-axis gyroscope) supports motion posture detection and step counting. The PCF85063 RTC connects to the batt via the AXP2101 for uninterrupted power supply. (Batt is not included)
  • Power Management and Abundant Interfaces---The AXP2101 power management system supports multiple output voltages, charging management, batt management, and lifespan optimization. It features an onboard 3.7V MX1.25 lithium batt charging/discharging interface. It includes a Type-C interface and programmable side buttons for KEY and BOOT. One I2C, one UART, and one USB pad are provided for easy external connection and debugging. (Batt is not included)
  • CNC Metal Chassis and Development Scenarios---The CNC unibody metal casing is robust and provides excellent heat dissipation. Suitable for AI voice dialogue intelligent agent prototype development and functional verification scenarios.

Cloudflare reported in 2025 that 37% of the top 10,000 domains it examined had a robots.txt file. Among robots.txt files it found among those top domains, 7.8% disallowed GPTBot and 5.6% included Google-Extended. These are Cloudflare’s scoped 2025 observations, not estimates of all websites or a measure of compliance.

How can a website tell which agent is making a request?

A plain user-agent string is a claim in a request, not strong proof of identity. Cloudflare describes Web Bot Auth as a framework for cryptographic bot identity, intended to help publishers distinguish identified bots from spoofed labels and set policies for them.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In Cloudflare’s documented Pay Per Crawl setup, a publisher creates an Ed25519 key pair, hosts a public-key directory, registers the bot, and uses HTTP Message Signatures on requests. This is Cloudflare’s implementation path; adoption of bot identity mechanisms is still evolving, and the available evidence does not establish how widely sites or agents use it.

Rank #4
ESP32-S3 1.28inch Double Eye Round LCD AIoT Development Board, Dual 1.28inch IPS Displays, Dual-Core 240MHz Processor, Supports Wi-Fi & Bluetooth 5 & AI Speech Interaction, Onboard DIY Connectors
  • This is an AIoT microcontroller development board based on ESP32-S3 with double eye LCD displays, designed for makers and electronics enthusiasts, supporting 2.4GHz Wi-Fi and Bluetooth BLE 5.
  • It integrates high-capacity Flash and PSRAM, onboard Dual 1.28inch LCD 240 × 240 resolution displays which can smoothly run GUI programs such as LVGL. Additionally, it also integrates a microphone, speaker header, Lithium battery recharge circuit, and reserves a TF card slot and DIY expansion connectors.
  • It is suitable for the quick development based on ESP32-S3 such as HMI (Human-Machine Interface), double eye robotic agents, and AI voice-interactive toys. Whether you want to build a robot that can "wink", create an intelligent IoT Interface, design touch-controlled games, or develop futuristic wearable devices, this board is an ideal choice.
  • Onboard ES8311 audio codec and ES7210 audio ADC chip, equipped with standard microphone and speaker header, Supports AI speech interaction. Allows access to online large model platforms such as ChatGPT, DeepSeek, Doubao, etc.
  • Onboard TF card slot for convenient local storage expansion, and supports the storing and reading of data, images, audio files, and more. Onboard Lithium battery recharge management module, reserved 3.7V Lithium battery power supply header. Onboard SH1.0 14PIN connector, adapting UART, I2C and some IO interfaces, for easy DIY customization.

Delegating access without disclosing identity

Cloudflare says it announced Private Access Control Tokens (PACT) with Mozilla, Google, Microsoft, and Shopify. The idea is that one site can vouch anonymously for a legitimate agent, which can then present a token at another site. The intended benefit is lower friction for legitimate access, but an anonymous token should not be mistaken for disclosure of the user’s identity, nor does it mean every site accepts the token.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Can websites charge AI crawlers?

Yes, in the sense that providers are testing mechanisms for paid access; there is no established universal price or standard fee. Cloudflare’s July 2025 Pay Per Crawl announcement described a private beta in which a publisher could allow, charge, or block a crawler. In its design, a site could respond with HTTP 402 Payment Required and pricing information, while the crawler supplied authenticated payment intent. Cloudflare said it acted as Merchant of Record for that flow. This was a vendor-specific beta announcement, not evidence that paid crawling is generally available across the web. See Cloudflare’s Pay Per Crawl announcement.

Cloudflare described a later direction as Pay Per Use: compensation tied to content creating value rather than solely to a fetch. Its 2026 announcement named Ceramic.ai, which Cloudflare said would pay when opted-in publisher content appeared in AI search results and return queries, citations, and ranking; it also said You.com would let agents pay on demand for premium content. Cloudflare named beehiiv among platform collaborators on creator access controls. These are descriptions from Cloudflare, not evidence of broad adoption or a standard commercial arrangement. See Cloudflare’s 2026 announcement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
seeed studio reSpeaker XVF3800 4-Mic Array with XIAO ESP32S3, Bare Board
  • Built for Custom Integration: Keep control of the enclosure, mounting and final device layout. The open-board format fits robots, kiosks, custom voice devices and embedded prototypes where flexible mechanical integration matters.
  • Onboard Voice Processing: XVF3800 performs AEC, beamforming, de-reverberation, DoA, VAD, AGC and noise suppression before audio reaches your application, helping reduce downstream audio preprocessing.
  • 360° Far-Field Voice Capture: Four MEMS microphones in a circular array support speech pickup from different directions at distances up to 5 m, so users do not need to speak toward one fixed microphone position.
  • XIAO ESP32S3 for Embedded Voice: The pre-soldered XIAO adds Wi-Fi, Bluetooth Low Energy and MCU-side control for connected voice interfaces, local wake-word projects and custom embedded applications.
  • Firmware Options: Ships with Standard I2S firmware for XIAO ESP32S3 and is not a USB audio device by default; switch to USB firmware for host audio or use dedicated 48 kHz HA I2S firmware for Home Assistant and ESPHome Voice; configurations are separate.

There are competing approaches. Arcep’s January 2026 report describes Akamai working with TollBit and Skyfire on content monetization, alongside proposals such as ai.txt, W3C text-and-data-mining work, and IETF efforts to signal intended uses including indexing, training, generation, and search. These initiatives show active experimentation, not a settled industry standard. Arcep’s overview is available in its January 2026 report.

Which access approach should a publisher choose?

There is no single best policy for every site. The choice depends on what the publisher wants to permit and whether it can reliably enforce the distinction. These are practical decision axes inferred from the documented controls and proposals, not a measured comparison of deployments.

Decision axis Questions for the site owner
Activity Is the request for search discovery, a real-time user task, or model training? Should those uses have separate rules?
Enforcement Is a published preference sufficient, or must the site technically block, authenticate, or conditionally admit requests?
Identity and delegation Can the site verify the bot, and can the agent demonstrate that it is acting with legitimate authorization without exposing unnecessary user information?
User friction and privacy Will a login, consent step, payment, or confirmation protect the site without making legitimate tasks impractical or collecting more data than needed?
Operating cost and freshness How much repeated fetching occurs, and can the site provide efficient updates or a more suitable interface?
Compensation Should access be free, paid per crawl, or compensated when downstream use creates value? Can the site verify that use and administer payment?

A graduated policy can make the distinction explicit: allow search indexing where it brings useful discovery, require verified identity for live interaction, and separately restrict or negotiate training use. That approach is only practical if the site’s infrastructure can enforce its rules; a robots.txt entry alone cannot do so.

What is still unsettled?

Web Bot Auth, PACT, paid crawling, and intended-use signaling are developing mechanisms, not universal admission credentials. The cited sources describe proposals, implementations, and named partners, but do not establish how widely these systems are deployed or whether a token or payment will be accepted by a particular site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Nor do these technical controls settle broader legal questions about scraping, copyright, or terms of service. Those questions depend on facts and jurisdiction; the mechanisms described here explain how access can be signaled, verified, denied, or priced, not what the law requires in a specific case.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.