Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ethical web scraping is a project-level practice, not a special legal status: check permission and the site’s rules, collect only what you need, keep your requests low-impact, protect people whose data may appear, and stop when access is restricted or causes harm. A page being publicly visible does not, by itself, settle whether collecting or using its contents is lawful.

Is web scraping legal if a page is public?

There is no reliable yes-or-no answer for every country, website, dataset, and purpose. Public visibility does not remove privacy obligations: an October 2024 joint statement by privacy regulators says publicly accessible personal information is subject to data-protection and privacy laws in most jurisdictions. Depending on the project, contract, copyright, database rights, computer-access laws, confidentiality, and site-specific restrictions may also matter.

For personal data, ask what purpose justifies collection, which lawful basis applies, what notice or other transparency is required, how accurate the information is, and how much can be minimized or deleted. The European Data Protection Board’s 8 July 2026 announcement about scraping for generative AI discusses purpose limitation, transparency, accuracy, data minimization, and GDPR lawful-basis requirements. If special-category data is processed under the GDPR, the EDPB summary says both an Article 6 lawful basis and an applicable Article 9(2) exception are needed. Public availability or a research purpose does not automatically create that exception.

The EDPB said its Guidelines 03/2026 had been adopted while open for consultation, with feedback due 30 October 2026. That is the status described in its 8 July 2026 announcement; check the current status and seek jurisdiction-specific advice for consequential projects. Neither the announcement nor the other sources establishes a universal rule that all scraping is legal or illegal.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does robots.txt give permission to scrape?

No. The Robots Exclusion Protocol is a crawler convention, not a permission grant or an access-control system. RFC 9309 states, “These rules are not a form of access authorization.” It also says that when a crawler successfully downloads a robots.txt file, it “MUST follow the parseable rules.” In practice, read the applicable rules and comply with them, but separately review the site’s terms, obtain permission where needed, and assess applicable law.

Check the robots.txt file at the root of the exact host, such as https://example.com/robots.txt, and match the rules for the user-agent you will send. Do not assume a file for one subdomain or protocol covers another. Google’s documentation describes how Google scopes its interpretation to the host, protocol, and port of the robots.txt URL; those are Google-specific implementation details, not a universal parser specification for every crawler. Read the IETF’s RFC 9309 and, if relevant to Google crawling, Google’s robots.txt documentation.

A robots.txt file is not a security barrier. A disallow rule should be treated as a clear crawler instruction, not something to work around. Likewise, a path absent from the file is not proof of permission. Do not bypass login, defeat a CAPTCHA, evade a block, or disguise a crawler’s identity.

How to plan a responsible scrape

1. Define the purpose and scope

Write down the decision or question the dataset must support before collecting anything. Specify the hosts and pages needed, the fields to retain, who could be affected, how long the data will be kept, and who will receive it. Exclude credentials, private areas, and sensitive or identifying fields unless you have specific permission and a defensible legal basis to process them. If a smaller dataset answers the question, collect the smaller one.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Check authorization and choose the access route

Review current terms and API conditions for the particular host, subdomain, protocol, and purpose. Prefer an official API or written permission when available. Privacy regulators note that APIs can give a host more control through credentials, logs, and monitoring. But an API key or contract does not by itself make otherwise unlawful personal-data processing lawful.

Keep a record of what you checked, when you checked it, the approved purpose, and any limits the site set. Recheck before a new crawl, after a significant site change, or if you change the project’s purpose or recipients.

3. Set conservative operating limits

Identify the crawler in a clear user-agent string, fetch only necessary pages, avoid parallel bursts, and cache responses where appropriate. There is no universal request rate that is safe for every site: set limits using the site’s published instructions and capacity, then monitor response times and errors. Back off or stop when latency rises, errors recur, access is blocked, or the site objects.

RFC 9309 says crawlers generally should not use a cached robots.txt for more than 24 hours unless the file is unreachable. That is a recommendation about how long to cache the robots file, not a recommended interval between page requests. The RFC also sets a minimum robots.txt parser limit of 500 KiB; that is a technical parser requirement, not an allowance to collect that volume of data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Protect people and the dataset

Names, contact details, account information, location, health, political views, and other sensitive details can create privacy risks. Decide whether privacy laws apply, document purpose and lawful basis where required, minimize fields and retention, and provide transparency when required. Restrict access to collected data and define a deletion schedule rather than keeping it indefinitely.

For projects relying on collected facts, preserve the source URL and collection timestamp, use reliable sources, and validate accuracy before relying on the data. The EDPB summary specifically discusses source reliability, timestamps, and accuracy validation in generative-AI contexts; applying those controls more broadly is a prudent data-governance practice, not a universal legal checklist.

5. Establish stop conditions

Stop if permission is withdrawn, the site adds restrictions, a block or objection appears, or the service shows signs of distress. Pause and review if the crawl unexpectedly exposes sensitive information or gathers more than intended. A one-time review of rules does not guarantee ongoing permission.

A cautious one-page Python example

This standard-library example checks the target’s root robots.txt for a clearly identified user-agent, refuses to proceed unless that file returns HTTP 200 and permits the exact URL, then fetches just that page and prints its title. It is intentionally conservative; it does not check terms, grant permission, determine a lawful basis, or make a site-wide crawl ethical. Use it only for a target you are authorized to access. Change the URL to an approved page and review the site’s current rules first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from html.parser import HTMLParser
from urllib.error import HTTPError, URLError
from urllib.parse import urlsplit
from urllib.request import Request, urlopen
from urllib.robotparser import RobotFileParser

TARGET = "https://example.com/"
USER_AGENT = "ExampleResearchBot/1.0 (contact: crawler@example.org)"
TIMEOUT_SECONDS = 15

class TitleParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_title = False
        self.parts = []

    def handle_starttag(self, tag, attrs):
        if tag.lower() == "title":
            self.in_title = True

    def handle_endtag(self, tag):
        if tag.lower() == "title":
            self.in_title = False

    def handle_data(self, data):
        if self.in_title:
            self.parts.append(data)

parts = urlsplit(TARGET)
if parts.scheme not in ("http", "https") or not parts.netloc:
    raise ValueError("TARGET must be an absolute HTTP or HTTPS URL")
robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"

robots_request = Request(robots_url, headers={"User-Agent": USER_AGENT})
try:
    with urlopen(robots_request, timeout=TIMEOUT_SECONDS) as response:
        if response.status != 200:
            raise RuntimeError(f"Robots file returned HTTP {response.status}; stopping")
        robots_text = response.read().decode("utf-8", errors="replace")
except (HTTPError, URLError, TimeoutError) as error:
    raise RuntimeError(f"Could not verify robots.txt; stopping: {error}") from error

rules = RobotFileParser()
rules.set_url(robots_url)
rules.parse(robots_text.splitlines())
if not rules.can_fetch(USER_AGENT, TARGET):
    raise RuntimeError("robots.txt does not allow this URL for this crawler")

page_request = Request(TARGET, headers={"User-Agent": USER_AGENT})
try:
    with urlopen(page_request, timeout=TIMEOUT_SECONDS) as response:
        if response.status != 200:
            raise RuntimeError(f"Page returned HTTP {response.status}; stopping")
        content_type = response.headers.get_content_type()
        if content_type != "text/html":
            raise RuntimeError(f"Expected HTML, received {content_type}; stopping")
        html = response.read(2_000_000).decode("utf-8", errors="replace")
except (HTTPError, URLError, TimeoutError) as error:
    raise RuntimeError(f"Page request failed; stopping: {error}") from error

parser = TitleParser()
parser.feed(html)
print("Page title:", " ".join(" ".join(parser.parts).split()))

The 2,000,000-byte read cap is a bound in this example, not an ethical quota or safe limit for other hosts. It does not implement site-specific rate instructions, handle every robots.txt edge case, or crawl links. For a real project, add only the behavior the permission and target rules allow; do not turn a failed check into a reason to evade it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What to do when a crawl fails or conditions change

  • Robots file unavailable or unreadable: the example stops rather than assuming access is allowed. Confirm the correct host and seek guidance or permission before proceeding.
  • Robots rules disallow the URL: do not fetch it with that crawler. Ask the site for permission or use an authorized API or another approved route.
  • HTTP errors, timeouts, or rising latency: stop or substantially back off; check whether the site is under load or has changed its rules. Do not add parallel retries to force completion.
  • 403, CAPTCHA, or explicit block: treat it as a restriction or objection, not a puzzle to defeat. Do not rotate identities or disguise traffic.
  • Unexpected personal or sensitive data: pause collection, restrict access to what was obtained, reassess the purpose and legal basis, and delete data that is not justified or needed.
  • Site terms or project purpose changed: repeat the authorization, privacy, and scope review before restarting. Permission for one purpose or route does not automatically extend to another.

When a screenshot is useful—and when it is not

A screenshot can preserve a visual record of an authorized public page, but it is not a substitute for permission to collect data, an API, or a privacy analysis. If a project needs structured facts, use an authorized structured access route when possible; a screenshot alone may omit hidden content, create accessibility and accuracy issues, and still capture personal information.

ScreenshotNeo is a website screenshot API and MCP server for developers. Its visual-capture options can help when the approved task is to capture a page as an image or PDF. It does not authorize access to a target site or make an otherwise impermissible collection ethical.

Or skip the browser setup

For an authorized visual capture, one GET request can return a screenshot or PDF. This cURL example saves a WebP screenshot of the example URL; replace it with a URL you are authorized to capture. See the ScreenshotNeo API documentation for parameters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
  • Cookie and consent banners are accepted like a visitor’s, and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing status in X-Page-Verdict and X-Billed headers.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.