The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Can I use browser automation on this website? There is no universal yes or no. Check the target service’s current Terms of Service, acceptable-use rules, API terms, machine-readable instructions and technical controls for the exact account, purpose and access method you plan to use. A page that loads successfully is not necessarily a page you are contractually allowed to automate.
Start with the service, account and jurisdiction
Write down the specific website, account, country or region governing the account, and what the automation will do. “Browser automation” can mean a user-directed agent retrieving one page, a crawler indexing millions of pages, a price monitor, an account-action bot or a training-data collector. Those activities can be treated differently by the same service.
- Target: record the exact domain, subdomain and logged-in area.
- Identity: note whether you use a personal account, an organization account, an API key or no account.
- Purpose: describe retrieval, search/indexing, training, monitoring, testing or account actions in one sentence.
- Location: identify the governing terms and jurisdiction shown for your account.
- Frequency: estimate requests, concurrency, schedule and data volume.
Terms and site controls change. Record the version or date you checked, any permission you received and the workflow covered by that permission. Recheck after changing domains, credentials, volume or purpose.
Read the documents in the right order
1. General Terms of Service and acceptable-use rules
Search the current terms for automated, bot, robot, scrape, crawl, data mining, copy, access, reverse engineer, circumvent, account sharing and rate limits. Look for both outright prohibitions and conditional permission, such as an approved API, written consent, a research exception or limits on commercial use.
#1 Best Overall
Google’s general terms, for example, prohibit automated access that violates machine-readable instructions on its pages. That is an example of one provider’s contract, not a rule that automatically governs every website.
2. API documentation and API-specific terms
If an API exists, use only the access method described in its documentation and read any additional API terms. Google’s API terms say users must not misrepresent or mask identity or attempt to circumvent documented limits. They also address scraping API content into permanent copies and retaining cached copies beyond permitted periods unless expressly allowed by the content owner or applicable law.
Check these details before writing code:
- Which endpoints and authentication methods are allowed.
- Requests per minute, daily quotas, concurrency and pagination limits.
- Whether caching, retention, backups and derived datasets are allowed.
- Redistribution, display, resale and training restrictions.
- Requirements to identify your application or user.
- Webhook, export and deletion obligations.
3. robots.txt and other machine-readable instructions
Read https://example.com/robots.txt and any crawler, AI or automation policy linked by the site. Treat the directives as an important signal, not a permission grant. Cloudflare’s documentation states that robots.txt compliance is voluntary and that the file does not technically prevent a crawler from requesting a URL. Conversely, a site’s failure to list a path does not prove that your contract permits automated access.
Honor applicable directives even when a technical workaround exists. Do not treat a disallowed path, a no-crawl directive or an owner’s published bot policy as something to evade.
Free tools Windows power users keep installed
One-click scans. No signup required.
4. Technical controls
Note login requirements, paywalls, CAPTCHAs, bot challenges, rate controls, geo-blocks and terms displayed inside the product. Passing a CAPTCHA, rotating identities or finding an unprotected endpoint demonstrates technical access only; it does not establish contractual authorization. If a control blocks the workflow, stop and seek an approved route.
Classify what your automation is doing
Cloudflare distinguishes Search (collecting or indexing content for later answers), Agent (automated activity in real time for a person, including browser-use agents) and Training (crawling to train or fine-tune a model). A single bot can have more than one behavior, and an owner can allow one category while blocking another.
| Behavior | Typical purpose | Questions to answer |
|---|---|---|
| Search | Indexing content for later retrieval | Are crawling, indexing, snippets and refresh frequency allowed? |
| Agent | Completing a real-time task for a person | Are logged-in actions, account data and delegated actions permitted? |
| Training | Building or fine-tuning a model | May content be copied, retained, transformed or used in datasets? |
| Other collection | Monitoring, testing, analytics or research | What volume, identity, storage and onward-use limits apply? |
Cloudflare describes a “Verified bot” as one that identifies itself honestly, behaves non-abusively, follows robots.txt and crawl directives, uses reasonable request rates and does not evade owner preferences. That framework is not a universal legal safe harbor. Cloudflare documentation updated in 2026 also described time-sensitive defaults for some new domains that were scheduled to block Training and Agent bots on pages displaying ads while leaving Search allowed. Verify the current documentation and the actual site configuration before relying on any such default.
Compare access methods before choosing one
| Access channel | What to verify | Common mistake |
|---|---|---|
| Documented API | Endpoint, credentials, quotas, caching, retention and redistribution terms | Assuming an API key permits unlimited copying |
| Normal browser session | Terms, account rules, consent notices, rate limits and allowed actions | Assuming human-visible pages are automatically bot-permitted |
| robots.txt or policy file | Disallowed paths, user-agent scope and update date | Treating a voluntary directive as either a license or a technical lock |
| Unauthenticated endpoint | Terms, copyright, privacy, rate and anti-abuse rules | Interpreting “publicly reachable” as “free to harvest” |
A practical pre-run checklist
- Identify the site, account, jurisdiction and exact URLs.
- Open the current Terms of Service and acceptable-use policy.
- Search for the automation and data-use terms listed above.
- Read API documentation and additional API terms if an API exists.
- Inspect robots.txt and linked crawler, AI or automation instructions.
- Describe the behavior: search, agent, training, monitoring or account action.
- Document request rate, concurrency, retention, sharing and deletion plans.
- Check login, paywall, CAPTCHA, geo and rate controls. Do not bypass them.
- Ask the owner for written permission when a clause is unclear or the use is material.
- Save the terms date, policy version, permission and workflow scope.
Common restrictions and what to do
“No bots,” “no scraping” or broad automated-access language
Stop treating the page as self-service. Look for an approved API, licensing process or written exception. Do not rely on a narrow interpretation based on how little data you collect.
API quota or rate-limit language
Implement the documented limit, backoff and pagination behavior. Never create extra keys, identities or parallel accounts to defeat a quota.
Account and identity restrictions
Use your own authorized account, identify the application as required and do not mask identity or share credentials when the terms prohibit it. For delegated actions, confirm that the account holder and service both allow the action.
Copying, caching or redistribution limits
Minimize stored content, enforce deletion dates and restrict downstream users. An API response may be viewable without granting permission to build a permanent mirror or commercial dataset.
CAPTCHA, bot challenge or paywall
Do not automate around the control. Use an official API, request access from the owner or redesign the workflow.
Unclear or conflicting signals
Terms, robots.txt and headers can point in different directions. Preserve the conflict, pause the run and seek clarification or qualified legal advice. The cited policies do not establish one universal answer for every site or jurisdiction.
Keep an audit record
For each project, store the terms URL and date, API documentation version, robots.txt copy, user agent, account owner, intended purpose, rate and retention settings, permissions and incident contacts. Recheck after a material change or on a schedule appropriate to the project. This record helps you prove what you believed you were authorized to do and lets you stop quickly when a policy changes.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If the job is to obtain a clean website image rather than interact with an account, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the page verdict and billing status.
One GET request returns PNG, JPEG, WebP or PDF. See the ScreenshotNeo documentation for parameters and policies.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minutecurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
It also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. Free accounts include 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create an account at ScreenshotNeo’s free sign-up.
Best Value
FAQ
Does robots.txt make automation illegal?
No single file answers that question. Robots.txt is a voluntary machine-readable instruction; combine it with the contract, API rules, owner controls and applicable law.
Is using an official API always allowed?
No. Use the documented method and comply with that API’s quotas, identity, retention, copying and redistribution terms.
Can a browser-use agent be treated differently from a search crawler?
Yes. Cloudflare’s categories separate real-time Agent activity from Search and Training, and site controls may handle them differently.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
What should I do when permission is ambiguous?
Pause the workflow, preserve the relevant policy text and request written permission or qualified advice rather than testing the boundary by bypassing a control.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

