Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sometimes—but there is no universal yes or no. Collecting pages anyone can view without logging in may present a different U.S. Computer Fraud and Abuse Act (CFAA) question from entering a protected account or bypassing a technical barrier. Even public-page scraping can raise contract, privacy, copyright, trespass, or other legal issues. In the EU, collecting personal data is processing under the GDPR and requires a lawful basis and compliance with its principles. Before scraping, assess how you access the site, what you collect, where the law applies, and what you will do with the data.

What determines whether web scraping is legal?

Web scraping is automated collection of information from websites. The fact that a script copies HTML does not decide whether the activity is lawful. Two projects that collect similar pages can have different risk profiles because one accesses public pages and the other logs in, one gathers business hours and the other gathers personal profiles, or one makes a few cached requests while the other imposes heavy load.

Assess the whole collection pipeline and its intended use. The relevant questions include:

  • Access: Is the content genuinely available to anyone, or does reaching it require an account, payment, or permission?
  • Barriers: Would collection bypass a CAPTCHA, IP block, paywall, login, or other technical control?
  • Data: Does the material identify or relate to a living person, or include sensitive personal information?
  • Rules and rights: What do the site’s terms, API rules, copyright notices, database rights, and robots.txt say?
  • Conduct and use: How much traffic will you generate, and will you publish, resell, profile people, or train an AI system with the results?
  • Jurisdiction: Where are you, the site operator, and the people whose data you collect—and which laws or contracts may apply?

No single factor is a complete answer. Public access may matter to a particular U.S. computer-access claim, for example, without settling a contract dispute or a GDPR question.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What does U.S. law say about scraping public websites?

The CFAA distinction: public pages versus protected areas

The federal CFAA restricts certain unauthorized access to computers. In its 2022 hiQ Labs v. LinkedIn opinion, the Ninth Circuit treated access to generally available, unauthenticated public pages differently from access to protected areas. It said hiQ had raised a serious question about whether the CFAA’s “without authorization” language covered collection of public LinkedIn profile pages, even after LinkedIn sent a cease-and-desist letter.

That opinion affirmed a preliminary injunction; it did not finally resolve every claim between the parties or make scraping categorically lawful. It is a regional appellate ruling, not a nationwide safe harbor. The factual distinction is important: viewing a page available to anyone is not the same access scenario as entering an account-only area or defeating a barrier.

DOJ charging policy is not blanket permission

The U.S. Department of Justice’s Justice Manual says prosecutors will not charge someone with “exceeding authorized access” solely for violating a contractual terms-of-service restriction on a generally available public website. The policy focuses on a narrower situation involving technically divided areas, access to some areas but not others, knowledge, and enforcement goals. It describes prosecutorial policy; it does not erase private contract claims or authorize circumvention.

Other U.S. claims can remain relevant

A weak or unavailable CFAA theory does not dispose of every possible claim. Site terms, registration terms, API licenses, copyright, trespass, state law, and—in relevant jurisdictions—database rights may still matter. The details of the site, your conduct, the material copied, and the use you make of it can change the analysis. Review applicable terms and notices before collection, rather than treating a public URL as blanket consent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does robots.txt make scraping illegal or give permission?

No. RFC 9309, the September 2022 standard for the Robots Exclusion Protocol, says robots.txt rules are crawler instructions that crawlers are requested to honor and expressly states they “are not a form of access authorization.” A disallow directive is not itself a statute, and an allow directive is not a license to collect data or override privacy, contract, or intellectual-property obligations.

Still, robots.txt is a meaningful operational and evidentiary signal. Record what it says at the time of collection and honor applicable directives as a risk-control practice. Do not treat its limits as an invitation to use a different route around an owner’s block. Authentication, paywalls, CAPTCHAs, IP blocks, and other technical controls are separate from robots.txt; do not bypass them.

Can you scrape personal data under GDPR?

Possibly, but public availability does not take personal data outside the GDPR. The European Commission defines personal data as information relating to an identified or identifiable living person. Processing is broad: it includes collection, recording, organization, storage, retrieval, consultation, use, and disclosure, whether automated or manual. The European Data Protection Board (EDPB) said on 8 July 2026 that GDPR applies when web scraping includes personal-data operations such as collection, storage, organization, and retrieval.

Identify the lawful basis and duties across the pipeline

Before collecting personal data, determine your role and identify a lawful basis under Article 6. Assess transparency duties—including whether an Article 14 exception is genuinely available—data-subject rights, international transfers, retention, security, and deletion. The GDPR principles apply throughout collection and reuse: lawfulness, fairness and transparency, purpose limitation, data minimization, accuracy, storage limitation, integrity and confidentiality, and accountability.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For accuracy and traceability, the EDPB emphasizes reliable sources, timestamps, and validation. Limit fields to what the purpose requires, set a retention period, and be able to correct or delete records. For high-risk work, document a data protection impact assessment (DPIA) and a necessity and balancing analysis.

Special-category data needs additional analysis

If a project involves special-category personal data, an Article 6 lawful basis alone is not enough: an applicable exception under Article 9(2) is also required. There is no blanket exemption for information that a person has posted publicly. If you cannot establish both requirements and address the other applicable duties, do not assume that public visibility makes collection acceptable.

Is scraping for generative AI allowed?

There is no general AI-training exception established here. The EDPB adopted Guidelines 03/2026 on web scraping in the context of generative AI in July 2026. As of 29 September 2026, the EDPB’s consultation page said feedback remained open through 30 October 2026. Treat the guidelines as current regulator guidance subject to consultation at that date, not as a new statute or a final rule that resolves every training project.

For an AI project, assess collection and downstream use separately. Ask whether the source permits the access, whether the corpus contains personal or special-category data, what lawful basis and transparency steps apply, how long the data and derived copies will be retained, and whether training, model release, or later reuse changes the purpose. A lawful route to access a page does not by itself settle the legal status of training on its contents.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do common scraping scenarios compare?

Scenario What raises or lowers risk Practical reading
Unauthenticated public pages, limited collection Generally public access is relevant to the CFAA distinction in the Ninth Circuit’s 2022 hiQ opinion. Terms, other claims, privacy, and use still matter. Not automatically unlawful or automatically permitted; review the other legal and operational issues.
Account-only content or use of a fake account Authentication and access to a technically separated area make the access question more serious. Do not assume the public-page reasoning applies; obtain permission or use an authorized API.
Bypassing a CAPTCHA, paywall, or IP block Attempts to defeat technical controls are distinct from visiting a generally available page. Stop rather than route around the control.
Public profiles containing personal data GDPR may apply to collection and later storage, organization, retrieval, or use; public posting does not remove those duties. Establish a lawful basis, address transparency and rights, and minimize data before proceeding.
Non-personal facts collected for AI training Access terms, copyright or database issues, technical barriers, load, and downstream use remain relevant; AI guidance is consultation-sensitive as of 29 September 2026. Do not treat “AI training” or “public page” as a legal permission by itself.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What should you do before a scrape?

  1. Define the project. Write down the purpose, jurisdictions, source sites, fields, expected volume, and retention period.
  2. Classify the fields. Separate non-personal data, personal data, and special-category data; exclude sensitive fields unless you have a defensible reason and legal basis to process them.
  3. Check site rules. Review terms, registration and login requirements, API rules, copyright or database notices, and robots.txt. Save a dated record of the applicable rules.
  4. Confirm authorized access. Do not build a workflow that depends on bypassing authentication, a CAPTCHA, a paywall, an IP block, or another technical barrier.
  5. Resolve privacy requirements. For personal data, identify your role and lawful basis, assess transparency and rights, minimize fields, document purpose and retention, and evaluate transfers and security.
  6. Reduce operational impact. Identify your crawler, use conservative rate limits, cache responses, retain source timestamps and provenance, and avoid unnecessary repeat requests.
  7. Respect objections and changes. Exclude sources that object, honor opt-outs and cease-and-desist communications, and reassess if the site changes access or its rules.
  8. Review downstream use. Reassess before resale, publication, profiling, or AI training; later use can raise issues beyond the act of fetching a page.
  9. Get advice when stakes are high. Seek jurisdiction-specific counsel for high-volume projects, sensitive data, cross-border collection, or a disputed demand.

What if a site sends a cease-and-desist?

Do not treat the letter as proof that every scrape was illegal, but do not ignore it or continue unchanged. Preserve the notice and relevant records, pause collection from the source while you assess the demand, and review the site’s terms, access method, data, volume, and downstream use. The hiQ opinion concerned a particular preliminary CFAA dispute; it does not decide a recipient’s contract, privacy, copyright, trespass, or other exposure. If the data includes personal information or the demand threatens litigation, consult a lawyer before responding or resuming.

Common compliance failures and how to correct them

  • “It loads in my browser, so it is fair game.” Public visibility is only one access fact. Recheck terms, privacy obligations, content rights, and intended use.
  • “Robots.txt permits it, so I have permission.” RFC 9309 says the protocol is not access authorization. Seek any license or consent the project actually requires.
  • “The data is public, so GDPR does not apply.” Public personal data can still be processed. Stop collection until you establish a lawful basis and address the relevant GDPR duties.
  • “A cease-and-desist cannot matter because of hiQ.” The Ninth Circuit decision was limited to a preliminary CFAA issue. Pause, preserve records, and assess other claims with appropriate counsel.
  • “We can solve a block by changing IPs or automating a login.” That changes the access facts and can sharply raise risk. Do not design around a technical barrier.

Or skip the browser setup

If your actual need is a visual record of a page rather than its underlying text or data, a screenshot API may fit better than building a browser-capture workflow. ScreenshotNeo is a website screenshot API and MCP server from Yorker Media, not legal permission to access a site or a substitute for privacy review. Its API accepts a URL and returns a PNG, JPEG, WebP, or PDF; clean-shot options accept cookie banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture, and each step can be turned off.

One cURL request:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

For Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

For Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options and response details. Bot checks, blank pages, timeouts, and failed loads are not billed; cache hits are also free, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Those operational features do not authorize a capture where the site’s rules or applicable law prohibit it. Visit ScreenshotNeo or sign up free for 1,000 screenshots a month with no card.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.