Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sometimes—but there is no universal yes-or-no answer. Web scraping’s legal risks depend on where the people and organizations involved are located, whether the pages are public or access-restricted, what information is collected, the site’s terms and technical controls, and how the results are used. Public visibility is not a general license to copy or reuse material. Conversely, a site’s terms-of-service restriction does not automatically make every scrape a crime. Those are separate questions, and privacy, copyright, database-right, contract, and anti-circumvention rules may also apply.

This guide explains the main distinctions, including what the U.S. hiQ litigation did—and did not—decide, why public personal data still raises privacy issues in the EU, and what to check before collecting data. It is general information, not legal advice for a particular project.

What makes web scraping legal or illegal?

“Web scraping” describes a method of collecting information from websites; it is not a single legal category. The same automated request can raise different issues depending on the access method, the material copied, the data subjects, the scale of collection, the intended use, and the governing law. A legal review should therefore separate the questions rather than treating “public” or “against the terms” as a complete answer.

Question What to examine
Access Can anyone view the page without an account, or does collection require login, defeating a technical measure, or continuing after access is blocked?
Terms and permissions Do site terms, an API agreement, a license, or another contract restrict automated collection or reuse?
Information collected Are you collecting personal data, factual information, expressive works such as text or images, or a substantial part of a database?
Purpose and reuse Will results be analyzed internally, republished, sold, used to contact people, or combined with other information?
Location Which countries’ laws may apply based on the operator, site, data subjects, and intended users?

These factors interact; none is a universal safe-or-unsafe switch. A page may be publicly viewable but contain personal data or copyrighted expression. A term may create contractual exposure without, by itself, resolving whether a criminal computer-access law was violated. A technical block can raise different concerns from a page that loads normally for any visitor.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is scraping public data legal?

Public access can matter, especially in the U.S. analysis of certain claims under the Computer Fraud and Abuse Act (CFAA), but it does not settle every legal question. Information visible without authentication is different from information behind an account or access control. Even for an openly available page, collection and reuse may still implicate privacy, copyright, database rights, contract, or other laws.

What the hiQ decisions say in the United States

The Ninth Circuit’s decisions in the hiQ litigation concerned publicly accessible LinkedIn data and a particular CFAA theory. In its 2019 appeal, the court discussed the distinction between information readily available to the general public and information kept confidential behind access restrictions. In 2022, it again addressed publicly accessible data under the CFAA while recognizing that other legal claims could remain.

Those decisions are not nationwide permission to scrape. They are decisions from one federal appellate circuit in a particular dispute; they do not answer every CFAA question elsewhere, approve scraping restricted accounts, or immunize conduct that may violate other laws. Do not read them as a blanket right to ignore technical restrictions, site terms, privacy duties, or intellectual-property rights.

Terms-of-service restrictions are a separate question

The U.S. Department of Justice’s Justice Manual, § 9-48.000, says prosecutors may not bring an “exceeds authorized access” CFAA prosecution solely on the theory that someone violated an access restriction in a contract or terms of service for a generally available internet service. That is federal prosecution policy, not a ruling that terms are unenforceable in private disputes. It does not eliminate possible contract claims, state computer-law claims, privacy claims, or other legal theories.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Justice Manual also states that it creates no enforceable right for a party in litigation with the United States. Treat a question about criminal charging policy as distinct from whether a particular contract applies or whether a private party could pursue a claim.

Can I scrape a website without permission?

There is no general rule that permission is always required for every automated request, or that permission is unnecessary whenever a page is public. The answer depends on the applicable law and facts. For example, a site’s ordinary public page, a signed-in account, and a page available only after overcoming an access-control measure are materially different situations.

Read the site’s terms, API rules, licenses, and relevant notices before collection. Consider whether the site offers an authorized API or another permitted route. A robots.txt file is a crawling signal; its presence or absence is not, by itself, a complete legal determination or a substitute for reviewing other obligations.

If a site blocks requests, sends a cease-and-desist letter, or disputes your access or reuse, do not assume that changing IP addresses or continuing under a different identity resolves the issue. Stop and reassess the legal basis and technical method, especially before resuming commercial or large-scale collection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I scrape personal data from public websites?

Public visibility does not remove personal data from privacy law. Under the EU General Data Protection Regulation (GDPR), personal data must be processed lawfully, fairly, and transparently, for specified and legitimate purposes, limited to what is necessary, kept accurate and no longer identifiable than necessary, and protected with appropriate security. Article 6 also requires at least one lawful basis for processing.

The GDPR states: “Personal data shall be processed lawfully, fairly and in a transparent manner in relation to the data subject (‘lawfulness, fairness and transparency’).” Whether GDPR applies, and what a particular controller must do, depends on the facts and the regulation’s scope. A public profile or directory entry does not by itself establish a lawful basis or answer questions about transparency, data-subject rights, or the purpose of collection.

Before collecting personal data, document the project’s purpose and assess whether the information is necessary for it. Consider how long it will be retained, how it will be secured, how people can exercise applicable rights, and whether special-category or otherwise sensitive information is involved. The analysis can become more demanding when data is combined, republished, used to make decisions about people, or collected across borders. Seek qualified privacy advice for a commercial project or sensitive data.

Does scraping copy copyrighted content?

It can. Extracting a fact is not the same as copying a page’s original wording, photographs, illustrations, or other protected expression. A project may collect factual fields while also storing or republishing expressive material. The intended use and applicable copyright law matter, so do not assume that calling a process “data extraction” resolves the copyright question.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Access and copyright are separate issues. A page’s public availability does not automatically authorize reuse of its contents, and a conclusion about computer access does not decide whether copying or distributing protected material is permitted. Assess the specific material and intended use, and get advice where the collection or republication is commercially important.

Can bypassing a technical block create a separate legal issue?

Potentially. The U.S. Copyright Office explains that Section 1201 of the Digital Millennium Copyright Act (DMCA) generally prohibits circumvention of technological measures used to control access to copyrighted works, subject to statutory provisions and exemptions. Whether a particular measure, work, act, or exception is covered is fact-dependent. Do not treat a CAPTCHA, login wall, or other control as merely an inconvenience to work around without considering the legal consequences.

The method matters alongside the outcome: a screenshot or dataset that appears innocuous does not establish that the way it was obtained was lawful. Conversely, not every technical restriction necessarily has the same legal status. Avoid broad conclusions without advice based on the relevant measure and circumstances.

What does Ryanair v PR Aviation mean for scraping in the EU?

The Court of Justice of the European Union’s Ryanair v PR Aviation judgment involved commercial extraction of flight data and Ryanair website terms that restricted screen scraping. The dispute addressed the relationship between those contractual conditions and EU database-right rules. It is a fact-specific decision, not a universal rule that all website terms are enforceable or that all commercial scraping is prohibited.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a new project, the actual terms, the nature and protection of the database, the method and scale of extraction, and applicable national law may matter. If you plan systematic or commercial extraction of a website’s database, do not rely on a one-line summary of that case as a substitute for reviewing the project.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical review before you scrape

  1. Map the jurisdictions. Identify where the site operator, your organization, data subjects, and intended users are located. Determine which laws may apply rather than assuming the law where your server sits is the only relevant one.
  2. Classify access. Record whether pages are available without authentication. Identify any login, access-control measure, block, or other restriction that the collection method would encounter. Do not bypass controls without legal review.
  3. Read the rules that govern the site. Check terms of service, API rules, licenses, and relevant notices. Treat robots.txt as a crawling signal, not a definitive legal verdict.
  4. Separate facts from protected material. Decide whether the project needs facts alone or also text, images, or other expressive content. Consider whether systematic extraction may involve a protected database.
  5. Assess personal data before collection. Document purpose, lawful basis where required, necessity, transparency, retention, security, and how rights requests will be handled. Escalate sensitive data or cross-border projects for qualified advice.
  6. Limit collection to the project’s needs. Use reasonable request rates and collect only what is necessary. A technical ability to retrieve more information does not establish a legal basis to do so.
  7. Set a stop-and-review trigger. Reassess if the site blocks you, challenges the activity, changes its terms, or you receive a demand letter. For commercial, sensitive, access-restricted, or disputed collection, consult a lawyer familiar with the relevant jurisdictions.

This checklist is a way to identify issues, not a legal clearance test. Laws differ by jurisdiction and can change, so verify current local requirements for your project.

Or skip the browser setup

If your goal is a visual record of a page rather than extracting its underlying data, you can request a screenshot without setting up a browser. ScreenshotNeo is a website screenshot API and MCP server for developers; it is not a legal permission to scrape or reuse a site’s content. You still need to consider the site’s rules and applicable law.

One GET request returns a screenshot or PDF. For example, this cURL request saves a WebP screenshot of Stripe:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for the request options. Before capture, it can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the page verdict and billing status. Its MCP server provides tools for AI agents to take screenshots, get page information, and capture PDFs.

The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Every feature is available on every plan. Learn more about ScreenshotNeo, or sign up for the free plan.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.