iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To see which bots and AI crawlers visit a React site, read the request logs at the layer that receives the traffic: your hosting provider, web server, or CDN. Filter the user-agent field for known bot names, then read each match with its timestamp, requested path, and HTTP status code. A React build does not include a logging console of its own, so the exact screens you use depend on where the site is deployed. A bot’s name in a log is a useful clue, but it is not proof of identity on its own.
Where crawler requests are recorded
A React app is delivered as built files, usually HTML, JavaScript, CSS, and assets, by whatever service answers requests for your domain. Every request a crawler makes reaches that service first, so the logs kept there show crawler visits whether or not the crawler runs your JavaScript. Three places are common:
- A hosting provider’s analytics or logs screen. Many static hosts and platform-as-a-service providers keep request logs in their dashboard. Look for a section named Logs, Requests, or Analytics in the project settings.
- A web server access log. If you run your own server, such as nginx or Apache, the access log records every request. Its location depends on the configuration, often under
/var/log/. - A CDN log or analytics view. If a CDN sits in front of the site, it sees the requests before your host does. Cloudflare, for example, lets you search logs by crawler user-agent and see requests, pages, and frequency for them.
Retention periods differ by provider and plan. Confirm how long your logs are kept before you need them, because a crawler pattern you want to investigate may already be gone.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Filter the logs for crawler names
- Find the layer that receives requests first. If a CDN is in front of the site, start with its logs. If not, use the host’s logs or the server’s access log.
- Filter the user-agent field for known tokens. Begin with the names in the table below. Your filter should be case-insensitive, because the same token can appear in different capitalizations in different tools.
- Keep the context for each hit. Record the timestamp, the requested path, the status code, and the full user-agent string together, so one record answers what was requested, when, and what happened.
- Classify the visit by purpose before you interpret it. Search, training-related crawling, and user-directed page retrieval are different activities, and the operator’s documentation determines which one a token represents.
- Mark the confidence level. Describe a hit as user-agent-matched unless you also checked the operator’s published verification method.
On a server that writes nginx’s default combined-format access log, a search like the following lists matching requests. Log locations and formats vary, so treat these as a pattern to adapt rather than a fixed recipe:
#1 Best Overall
- Hardware Controller with Professional Network Management-Centralized management for up to 100 Omada devices including Omada access points, Omada Security Gateways and Jetstream switches.
- Premium Hardware Design-Industry-leading flexible Rackmount/Desktop design with a powerful chipset, durable metal casing, 2 fast ethernet ports and 1 USB 2.0 port for auto backup.
- Dual power selection-Support PoE (802.3af/802.3at) and micro USB for flexible installations.
- Easy Network Monitor & Maintenance-The easy-to-use dashboard makes it simple to see your real-time network status and improve network maintenance for peace of mind.
- Cloud Access with No License Fee-Enjoy cloud service with no license fee with the use of OC200. Remote Cloud access and Omada app brings centralized cloud management of the whole network from different sites—all controlled from a single interface anywhere, anytime.
grep -iE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Googlebot' /var/log/nginx/access.log
In that format, the requested path is the seventh space-separated field and the status code is the ninth. To count which paths a token requested most often:
grep -iE 'GPTBot' /var/log/nginx/access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -20
To see the status-code distribution for the same token, replace $7 with $9. Each command only shows requests your server actually logged, so requests blocked at a CDN before reaching your server will not appear in this view.
Which bot names to recognize
The table lists tokens named in Cloudflare’s bot reference, which was last updated 2026-04-23, and in the operators’ own documentation, checked in October 2026. Cloudflare describes its list as a selection rather than a complete directory, and it points to Cloudflare Radar for the current directory. Operators also add and rename tokens, so check the live pages before relying on any single name.
Recommended Free Tools
Rank #2
- Automatic Router Rebooter / Reset - Stop manually restarting your router! Automate the process to ensure highly reliable internet connection uptime
- Constantly Monitors Router and/or Modem Internet Health. Keep Connect provides 24/7/365 protection to ensure that your smart home and connected devices are always online and available.
- Notifications - Free Texts or Emails from Keep Connect notifying you of detected eventsif you choose to enter your phone number/email. You may also choose No Notifications.
- Perfect for Smart Home Reliability - Schedule Periodic Resets to keep your connection fresh and fast.
- Premium Cloud Services App Available (iOS App Store and Google Play Store) - Our Premium Keep Connect Cloud Services platform allows using our Online/Mobile App to monitor many locations in one place as well. Cloud Services allows remote management of devices at all locations as well as heartbeat monitoring of your Keep Connects to notify you in the event of an ISP internet outage at one of your sites.
| Token | Operator | Stated purpose |
|---|---|---|
GPTBot |
OpenAI | Crawling that may be used for training OpenAI’s generative AI foundation models, according to OpenAI’s crawler documentation (checked 2026-10-07). |
OAI-SearchBot |
OpenAI | Crawling for ChatGPT search results. OpenAI states that this setting is independent of GPTBot. |
ChatGPT-User |
OpenAI | User-initiated page access. Cloudflare classifies it as an AI assistant. |
ClaudeBot |
Anthropic | Listed by Anthropic’s help center (dated 2026-04-07) as one of three Anthropic crawlers. Confirm its stated purpose on that page. |
Claude-SearchBot |
Anthropic | Listed separately from ClaudeBot and Claude-User in Anthropic’s help center. Confirm its stated purpose on that page. |
Claude-User |
Anthropic | Listed separately from ClaudeBot and Claude-SearchBot in Anthropic’s help center. Confirm its stated purpose on that page. |
PerplexityBot |
Perplexity | AI search, according to Cloudflare’s bot reference. |
Googlebot |
Ordinary search crawler. Google documents its subtypes and identifies them through the HTTP user-agent header (Google Search Central, “What Is Googlebot,” checked 2026-10-07). |
Googlebot deserves a specific caution. Crawling and indexing are separate steps. A page that Googlebot did not request in your logs has not necessarily been dropped from search results, and a logged request does not guarantee that the page was indexed.
Read each hit with its path and status code
A matched request is only the start. Its status code tells you what the server or CDN did with it:
- 200 means the logged layer returned a successful response, so the crawler received the content.
- 403 or 429 usually means the request was refused or rate-limited, which is a different outcome from a page fetch. Check the CDN’s rules or the host’s settings to see which one applied.
- 404 means the path does not exist. Repeated 404s for paths you never published often point to scanners or stale links rather than a site problem.
- 5xx responses mean the server failed to answer. These are worth investigating on your own site, because they can affect the crawler’s next visit.
Paths matter as much as counts. A crawler requesting your home page and blog posts is behaving differently from one requesting your API routes or login pages, and that difference is easy to miss if you only count hits.
Rank #3
- (10/100/1G) Gigabit Bypass network tap / sniffer equivalent to port mirror on a switch.
- The two monitor/sniff ports are isolated from the network being monitored.
- Automatic bypass of device on power fail.
- Power-over-Ethernet (POE) pass-through. Rated at .75A max at 57vdc
- 5v power through USB3 port or 5v wall transformer (or both). ~500ma consumption.
How sure can you be that a bot is who it says it is?
A user-agent string is a clue, not verification
Any client can send any user-agent string, and some bots do not send an identifying header at all. Cloudflare notes that such bots may need other signals to identify. For that reason, a log line that says GPTBot establishes that the request claimed to be GPTBot. It does not establish that OpenAI sent it.
Verification signals you can check
- Source IP address. OpenAI publishes IP ranges for its documented crawlers. Compare the client IP in your logs with the current published list. If a CDN or proxy sits in front of your server, the logged address may belong to the proxy rather than the visitor, so confirm which address your log records.
- Operator-published verification. Google documents how Googlebot is identified through its user-agent header, and you should apply the verification method each operator documents before calling a hit verified.
- Consistency over time. Repeated requests from a single address range that matches a documented operator are stronger evidence than one line in a log.
Until you have checked a hit this way, describe it as user-agent-matched in any notes or reports.
What a visit does and does not tell you
A crawler request is not automatically training data collection, and it is not automatically a search index update either. OpenAI’s documentation describes its three crawlers as separate settings, and Cloudflare’s bot categories separate search, agent, and training behavior. Use the operator’s current documentation to classify each token, and avoid calling every AI-related request a training crawler. A user-initiated request such as ChatGPT-User, for example, reflects a person asking for a page, which is a different situation from a scheduled crawl.
Rank #4
- NEVER MANUALLY REBOOT YOUR ROUTER AGAIN – The ConnectSense Rebooter Pro plugs between your modem or router and the wall outlet, automatically detecting lost internet connectivity across up to 5 network targets and power cycling your equipment instantly — keeping your home, office, or remote location always online 24/7.
- SCHEDULED & AUTOMATIC REBOOTS – Set up to 10 custom reboot schedules to proactively clear memory leaks, prevent slowdowns, and keep your connection fresh — even before problems occur. Perfect for smart homes, security cameras, smart locks, thermostats, and any device that depends on a stable internet connection.
- REMOTE CONTROL FROM ANYWHERE – Trigger a manual reboot anytime from the free ConnectSense app (iOS & Android) or directly from your home network. Whether you're traveling, at work, or managing a vacation rental or remote office, you stay in control of your network without needing to be on-site.
- AUTOMATIC POWER OUTAGE RECOVERY – When the power goes out, the Rebooter Pro automatically restores and reboots your networking equipment once power returns, eliminating downtime and the need for manual intervention. Ideal for unattended locations, rental properties, and small business networks.
- INTEGRATOR & PRO-GRADE FEATURES – The only router rebooter with a built-in local HTTPS API, giving IT professionals, smart home integrators, and power users advanced automation, monitoring, and remote management capabilities — no cloud subscription required for local control.
robots.txt states preferences; it does not enforce them
A robots.txt file at the root of your domain tells crawlers which paths you prefer they avoid. Compliant crawlers read it and follow it, but the file does not stop a request from reaching your server. Cloudflare states the point directly in its AI-crawler guidance: “Robots.txt is not binding — following it is more of a courtesy than anything else.” Attribute that statement to Cloudflare.
OpenAI’s documentation explains that each of its crawler settings works independently. Its published example reads: “Each setting is independent of the others – for example, a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training OpenAI’s generative AI foundation models.” Anthropic’s help center says its bots honor standard robots directives. That is Anthropic’s own statement about its bots, not a guarantee about every crawler on the web.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Two practical points follow. First, confirm that the file you edited is the one served at your domain. A React project often keeps static files in a public folder, and a file that never reaches the deployed build will not be read. Second, if you need technical enforcement, use controls at the host or CDN, such as firewall or bot-management rules. Blocking a request at the CDN also changes what appears in your logs, so keep that in mind when you compare counts over time.
Best Value
- [UPGRADED NanoVNA-H] New HW Version V3.7. It is upgradeable as new firmware is developed. With MicroSD card port now can have the measurement data or the screenshots saved in the it at anytime. Added battery circuit management, more secure. Redesigned PCB, you can connect to mobile phone with Type C-Type C cable (original PCB needs OTG cable), see a clear HD image on your phone. Added a ABS case, which is protective and dust-proof. Disply: 2.8 inch TFT (320 x240).
- [IMPROVED FREQUENCY ALGORITHM] The improved frequency algorithm can use the odd harmonic extension of si5351 to support the measurement frequency up to 1.5GHz. The 9KHz-300MHz frequency range of the si5351 direct output provides better than 70dB dynamic, The extended 300M-900MHz band provides better than 60dB of dynamics, and the 900M-1.5GHz band is better than 40dB of dynamics.
- [MULTIPLE FUNCTIONS] The default firmware main function is used for antenna performance measurement. The TX/RX method can measure the complete S11 and S21 parameters. If you need to obtain S12 and S22, you need to manually replace the transceiver port wiring. The CH0 output level is increased to 0dBm when using the fundamental wave, resulting in more accurate reflection measurement.
- [SUPPORT ANDROID PHONE & PC SOFTSARE CONTROL] Designed a practical and simple control application on PC, you can download touchstone(SNP) files for radio design and simulation software. There is a PC interface that adds functionality and lets you work interactively on a bigger screen. Supports time domain analysis function (TDR). Compatible with most Android mobile phones, convenient for connecting to mobile phones. Support Windows Computer Control.
- [STRONG AND SECURE POWER SUPPLY] This VNA is battery powered or USB powered. Built in 650mAh battery, could work for 2 hours continuously. For longer measurement time, kindly connect an external power source. The product interface displays battery usage, providing a clear understanding of the power status.
Optional: Cloudflare AI Crawl Control
If your domain already runs through Cloudflare, AI Crawl Control gives a managed view of crawler activity. Cloudflare’s documentation (last updated 2026-04-23) describes analytics for crawler request volume, allowed requests, status-code distribution, popular paths, and operators, along with filters and individual crawler controls. Those views let you check the same questions covered above without writing log filters.
Two qualifications apply. Referral analytics are available on paid Cloudflare plans, according to the same documentation, so check your plan before you plan around that feature. Second, the tool is optional. You do not need it to inspect logs, and a site that does not use Cloudflare cannot use it at all.
A simple way to choose between the two approaches:
- Use direct log filtering if you already have access to your host’s or server’s logs and want full control over the filters.
- Use AI Crawl Control if Cloudflare already sits in your traffic path and you want status, path, and operator views without building your own queries.
- Use both if you need the Cloudflare view for daily monitoring and the raw logs for verification of specific hits.
Checks throughout this guide reflect Cloudflare, OpenAI, Anthropic, and Google documentation as checked on 2026-10-07. Bot names, IP ranges, and plan terms change, so confirm them on the operator’s current pages before acting on them.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

