Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Web scraping is not automatically legal or illegal just because the information is publicly visible. The answer depends on where the scraping occurs, how the site is accessed, whether applicable terms restrict it, what data is collected, and how that data is used. Website operators can detect and block scraping, but a technical block neither settles the legal question nor guarantees that collection will stop.

When is web scraping legal?

There is no single rule that answers this for every website or country. Public visibility is one relevant fact, not a complete legal analysis. Access restrictions, applicable terms, the nature of the data, the scraper’s purpose and later use, and the governing jurisdiction can all matter.

The available examples below are deliberately limited: a U.S. appellate decision addresses one question under one federal law, while French regulator guidance addresses safeguards for scraping personal data. Neither resolves every legal issue. In particular, these sources do not establish a comprehensive answer on copyright, database rights, or the laws of every jurisdiction.

What the U.S. hiQ decision does—and does not—say

In hiQ Labs, Inc. v. LinkedIn Corporation, the Ninth Circuit’s April 18, 2022 opinion considered whether hiQ’s continued collection of LinkedIn profiles visible to anyone with a web browser was access “without authorization” under the Computer Fraud and Abuse Act (CFAA). The court’s holding was about that CFAA question and those publicly viewable profiles; it is not a general ruling that scraping is lawful. Read the Ninth Circuit opinion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The court also noted that other remedies may remain available even if the CFAA does not apply: “Entities that view themselves as victims of data scraping are not without resort, even if the CFAA does not apply: state law trespass to chattels claims may still be available.” That passage identifies a possible claim, not a finding that every scraping activity creates liability.

Why terms of service are a separate issue

The LinkedIn User Agreement quoted in the appellate opinion prohibited scraping and copying profiles and using automated methods to access the service. A conclusion about the CFAA does not decide whether a contract claim applies. Whether terms bind a particular collector depends on facts such as whether the collector assented to them and what conduct occurred. The later district-court record describes a breach-of-contract claim and disputes about defenses. See the Northern District of California record.

Why personal data changes the analysis

France’s data-protection authority, CNIL, says in guidance published June 19, 2025, that scraping personal data generally relies on legitimate interests and should include additional safeguards to limit effects on people’s rights and freedoms. Its recommendations address reasonable expectations, sensitive data, default exclusions for some sites containing particularly intrusive or sensitive information, and ways to make it easier for people to exercise a prior right to object. Read CNIL’s guidance.

CNIL states: “La collecte des données accessibles en ligne par moissonnage (web scraping) doit être accompagnée de mesures visant à garantir les droits des personnes concernées.” In English: online data collection by scraping must be accompanied by measures to safeguard the rights of the people concerned. This is French regulator guidance, not a complete answer to every GDPR question or the law in other countries. Information being viewable online does not by itself determine whether it may be collected and reused for a particular purpose.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does it matter whether pages are public or restricted?

Yes. The hiQ opinion concerned profile pages visible to anyone with a web browser; it should not be stretched to cover logged-in, restricted, or otherwise gated areas. Consider the circumstances together rather than treating “public” as a permission slip:

  • Access: Is the material openly viewable, or does access require an account or other gate?
  • Terms: Are there terms that restrict automated access, and did the collector accept terms that apply to the conduct?
  • Data: Is the collection limited to ordinary public information, or does it include personal, sensitive, or highly intrusive information?
  • Purpose and reuse: Why is the data being collected, and what will happen to it afterward?
  • Jurisdiction: Which country’s rules and which court or regulator’s guidance are relevant?

These factors help identify what needs review; they do not produce a universal yes-or-no result. For a commercial project involving personal data, jurisdiction-specific legal review can help assess the applicable rules and safeguards.

Can a website prevent web scraping?

Operators can deploy technical measures to detect, monitor, and block scraping activity. The LinkedIn district-court record describes the company sending a cease-and-desist letter and implementing such measures. That supports the practical point that blocking is possible; it does not establish that any particular control will stop all collection. The cited record does not compare products, engineering approaches, or their effectiveness.

For an operator, the prevention goal matters: reducing automated access, managing load, protecting sensitive areas, and documenting a site policy are different objectives. Published rules and access controls can be combined with monitoring and blocking, with legal review of the site’s terms and data practices. These are operational measures, not guarantees or legal conclusions. A scraper’s ability to evade a control would not itself establish permission to collect the data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Does robots.txt make scraping illegal?

Do not treat robots.txt as a universal legal ruling. The litigation record mentions it, but the sources cited here do not establish its legal effect across jurisdictions. A site can use it as an instruction to automated crawlers and an operational signal; whether ignoring it has legal consequences depends on the relevant facts and law. Check the site’s terms and the applicable jurisdiction rather than relying on robots.txt alone.

A practical decision check before collecting data

  1. Identify the jurisdiction. Determine where the collector operates, where the site or affected people are located, and which legal rules may apply. The hiQ opinion is a Ninth Circuit decision; CNIL’s guidance is from France.
  2. Map access and terms. Record whether the target pages are public or gated, review applicable site terms, and assess whether the collector accepted restrictions on automated access.
  3. Classify the information. Determine whether the material includes personal or sensitive data and whether the collection could be particularly intrusive. CNIL’s guidance calls for safeguards when scraping personal data and discusses excluding some sensitive or intrusive sources.
  4. Specify purpose and reuse. Document why data is needed, what will be retained, and how it will be used. Public availability alone does not answer whether a particular collection and reuse are appropriate.
  5. Get focused legal review where needed. The cited authorities do not settle copyright, database rights, or law outside their stated scope. A project with meaningful legal exposure needs analysis for its jurisdiction and facts.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.