To see which bots and AI crawlers visit a Next.js site, log incoming request metadata at the server boundary—especially the raw User-Agent—and classify recognizable bots separately. Next.js can identify requests matching bots its helper knows about; that label is a clue, not proof of who sent the request. For Vercel-hosted sites, its managed AI-bot ruleset is a separate option for observing known crawlers.
What to log for each request
Keep the original request signal alongside any derived bot label. A practical record includes:
- UTC timestamp and HTTP method.
- Normalized path, without query values that might contain secrets.
- Response status and a provider request identifier, when available.
- Raw User-Agent, stored as received.
- Classification and provenance: for example, whether Next.js recognized a bot, a platform-provided bot family, and which system supplied the label.
Do not log credentials, cookies, authorization headers, or personal data you do not need. Send records to a controlled log destination and set retention to suit the operational question. A request count is not a count of unique visitors, and a crawler request does not by itself show that a page was indexed.
Choose where to observe requests
Middleware or Proxy for request-level logging
Next.js describes Middleware as server-side code that runs before a request completes and identifies logging as one use for custom server-side logic. In newer conventions, request interception may use Proxy; naming and framework conventions vary by version, so follow the documentation for the version deployed. A request-intercepting boundary is useful when you need to observe traffic before route rendering.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute#1 Best Overall
Use a narrow matcher so the code does not run on routes or assets you do not need to inspect. Middleware can be invoked across routes, so matcher scope affects both coverage and overhead. Keep logging lightweight and avoid turning every request into a large payload or a blocking operation. See the Next.js Middleware documentation for the framework’s behavior and configuration.
App Router Server Components for route-specific inspection
In an App Router Server Component, current Next.js documentation makes headers() asynchronous. Read the header with (await headers()).get('user-agent'). The returned headers are read-only. Using this Dynamic API makes the route dynamic, so this is suitable when the route itself needs request information, but it is not the same as a centralized log of requests before rendering.
Rank #2
Consult the Next.js headers() API reference for the deployed framework version.
Classify bots without treating the label as identity
When a NextRequest is available, Next.js provides the userAgent(request) helper. Its isBot property indicates whether the request matches a known bot in the helper’s recognition logic; the helper also parses browser and device information. See the Next.js userAgent API reference.
Rank #3
The helper works from request information, including the User-Agent string. That string is supplied by the requester and can be imitated. A recognized name such as Googlebot or an AI crawler is therefore useful for triage, not cryptographic verification of the operator or network origin. Nor does a false result establish that a request is human: the crawler may be new or simply unrecognized by the helper.
Next.js learning material describes crawlers as identifying themselves with custom User-Agent strings, and its metadata documentation describes HTML-limited bot handling based on the incoming User-Agent. Those framework behaviors do not verify crawler IP ownership. Keep the raw header so you can audit or revise a classification rather than relying only on a boolean.
Rank #4
Vercel-hosted sites: use managed AI-bot controls as another signal
Vercel documents an AI bots managed ruleset that identifies known AI crawlers and supports log or deny actions, using a maintained list. It is a platform feature, separate from framework-neutral Next.js logging; its availability and behavior apply to Vercel-hosted sites. See Vercel Bot Management documentation for the current directory and configuration.
If the goal is to find out who is visiting, begin in log or observation mode. Denying traffic changes what those crawlers can access and should follow an explicit policy, not be an automatic response to seeing a bot name. Vercel’s bot observability guidance also describes firewall visibility for IP, User-Agent, and request counts, as well as runtime logs and external telemetry as investigative inputs. It does not, on its own, establish setup steps or equivalent AI-bot classification for each external service; consult the relevant provider’s current documentation before choosing one.
Framework logging or managed controls?
These approaches answer different questions, and the available documentation does not establish a complete feature or price comparison.
Quick Recap
| Consideration | Next.js request logging | Vercel AI-bot managed ruleset |
|---|---|---|
| Bot coverage | userAgent(request) recognizes bots known to its helper; the source does not promise coverage of every new AI crawler. Next.js API reference |
Identifies known AI crawlers using a maintained list. Vercel documentation |
| Raw request context | Request APIs provide the User-Agent; what else you record depends on your implementation and hosting request context. Next.js headers reference | Vercel observability guidance describes visibility into IP, User-Agent, and request counts. Confirm which fields are available in the specific view you use. Vercel guidance |
| Observe before enforcement | You control what your logger records and can use labels without blocking requests. | Rules support log or deny actions; select log when observing rather than enforcing. Vercel documentation |
| Retention, export, querying, overhead, portability, plan limits, privacy controls | Depends on your implementation, log destination, and hosting environment; not specified as a complete comparison in the cited framework documentation. | Not stated as a complete comparison in the cited Vercel bot-management documentation; check current provider documentation and account details. |
A practical workflow
- Choose the observation boundary. Use a narrowly matched Middleware or Proxy for requests to inspect before route rendering; use
headers()when a specific App Router route needs the User-Agent. - Record raw evidence and context. Capture timestamp, method, normalized path, status, raw User-Agent, and provider request ID where available. Exclude query values and sensitive headers unless there is a justified need.
- Add a clearly sourced label. Store whether Next.js recognized a bot or whether a hosting rule identified a bot family. Preserve the raw User-Agent and note the source of each classification.
- Observe a baseline. Review names, paths, statuses, request rates, and time windows. Distinguish repeated requests from unique visitors, and requests from successful indexing.
- Set a policy only after review. Decide whether to allow, rate-limit, challenge, or deny selected traffic. For Vercel, check the current ruleset directory and mode behavior before enforcing a rule.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

