Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWeb crawling discovers and requests pages, usually by following links; web scraping extracts selected information from pages, feeds, or APIs. A crawler can collect pages without extracting structured fields, and a scraper can extract data from a single page without crawling a site. If you do either, prefer an official API or permissioned feed, respect access boundaries and site rules, and build in conservative request limits and a way to stop.
What is the difference between web crawling and web scraping?
A web crawler is an automated client that discovers and retrieves resources. It may start with a list of URLs, then follow links to find additional pages. Search engines use this recursive traversal to discover content for indexing. The IETF’s RFC 9309 describes crawlers as automated clients.
A web scraper focuses on extracting selected fields or content—such as page titles, product descriptions, or publication dates—from pages, feeds, or APIs. A scraper may process pages found by a crawler, but it may also work from a supplied list of URLs or one page.
| Question | Crawling | Scraping |
|---|---|---|
| Main job | Discover and retrieve resources, often by following links. | Extract selected data from a known source. |
| Typical output | A set of URLs or retrieved pages. | Structured fields or content. |
| Can it happen alone? | Yes. A crawler can discover pages without parsing their content into fields. | Yes. A scraper can process one supplied URL or an API response without discovering more pages. |
In practice, a production pipeline may crawl a permitted part of a site, scrape only the fields it needs, and store the source URL and retrieval time alongside each result.
#1 Best Overall
When should you use an API, a crawler, or a scraper?
Choose the least invasive and most stable source that satisfies the task. The right choice depends on permission, whether data is public or authenticated, whether the work is one-off or recurring, and whether the relevant content is present in static HTML or rendered by JavaScript.
| Approach | Best fit | Trade-offs to check |
|---|---|---|
| Official API, export, or permissioned feed | Recurring access to data the provider makes available through a defined interface. | Check the permitted use, fields, authentication, rate limits, stability, and retention terms. It is usually easier to govern and less brittle than extracting changing HTML. |
| HTML extraction from a known URL list | A limited, permitted task where the needed information is exposed in page markup and no suitable API or feed exists. | Markup changes can break parsers. Confirm the terms and access boundaries, minimize requests, and monitor for layout changes. |
| Crawling plus extraction | A permitted recurring task that requires discovering pages as well as collecting specific fields. | Link discovery can expand quickly. Scope the hosts and paths, set request and storage limits, and make stopping straightforward. |
| JavaScript-rendered page capture | When relevant page content is not available in the initial HTML and you have permission to retrieve it. | Rendering adds browser setup, time, and resource use. It does not grant access to restricted content or make extraction more stable by itself. |
For a one-off research question, a manual review or provider export may be simpler than building a crawler. For a recurring production job, compare self-hosted tooling with managed infrastructure on permission controls, rate control, observability, data protection, and maintenance—not just setup effort or request volume.
What does robots.txt do—and what does it not do?
A site publishes crawler instructions in a robots.txt file at the top level of a particular host, for example https://example.com/robots.txt. Under RFC 9309, a crawler reads the matching user-agent group and applies the most specific applicable allow or disallow path rule. The rules are requested crawler behavior; as the RFC states, “These rules are not a form of access authorization.”
- It guides crawler behavior. Treat a matching disallow rule as an instruction not to crawl that path.
- It is not a login or permission system. A robots file does not authorize access to a page, and it does not make bypassing authentication or technical controls acceptable.
- It does not reliably keep a page out of search results. Google Search Central describes robots.txt as a way to tell search engine crawlers which URLs they can access and says it is mainly for managing crawl traffic. For search exclusion, use
noindexor authentication as appropriate; robots.txt is not a reliable de-indexing control. - Its scope is host-specific. Rules apply to the relevant protocol, host, and port. Do not assume a file on one hostname governs a subdomain or a different protocol or port.
RFC 9309 distinguishes an unavailable robots.txt response from an unreachable server error. Decide how your crawler handles each condition rather than treating every fetch failure as permission to proceed. The RFC recommends conservative caching—generally no more than 24 hours unless the server is unreachable. Record the file and the time you used it so an operator can later understand the crawl decision.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteIs web scraping legal?
There is no universal yes-or-no answer. The legal analysis can depend on jurisdiction, what the data contains, whether the material is public or behind authentication, the site’s terms and notices, and exactly how collection and reuse occur. Public availability alone does not erase potential obligations involving contracts, copyright, privacy, or other legal claims.
In its 2022 opinion in hiQ Labs v. LinkedIn, the Ninth Circuit addressed a preliminary injunction involving public LinkedIn profiles. On that record, the court treated access to public sites as unlikely to be “without authorization” under the U.S. Computer Fraud and Abuse Act (CFAA). That limited conclusion is not a general license to scrape: the opinion also recognized that claims such as trespass to chattels, copyright, misappropriation, unjust enrichment, conversion, contract, and privacy claims may remain available.
Rank #3
Before collecting, identify the relevant jurisdiction and purpose, check the site’s terms, notices, authentication boundaries, and opt-out instructions, and assess the data’s sensitivity and intended use. If the task involves personal data, restricted content, or material you plan to publish or commercialize, obtain advice for the specific circumstances rather than relying on a broad rule about public pages.
How do you crawl or scrape a site responsibly?
- Define the job. Write down the purpose, fields, relevant geography, retention period, and lawful basis. Collect only what the task needs.
- Look for a supported source first. Prefer an official API, data export, or permissioned feed when one is available and suitable.
- Check scope and instructions. Fetch and parse the target host’s robots.txt; record the exact file and timestamp. Read terms, notices, authentication boundaries, and opt-out instructions.
- Do not cross access boundaries. Do not bypass logins, paywalls, CAPTCHAs, or technical access controls. Stop if the owner asks you to stop.
- Identify the client. Use a stable user-agent and, where appropriate, a contact address so site operators can identify the automated traffic.
- Control load. Keep concurrency low; use backoff, caching, and conditional requests where supported. Provide a kill switch. Stop on repeated 403, 429, or 5xx responses rather than escalating requests.
- Minimize and protect data. Keep only necessary fields, protect personal information, retain source URLs and retrieval timestamps, and provide a way to handle deletion or correction where applicable.
- Maintain and audit the pipeline. Test parsers against layout changes, monitor error rates, and keep an audit trail of permissions and decisions.
A small Python example for a permitted, public crawl
This example checks robots.txt, stays on one host, follows only links the policy allows for its declared user-agent, limits the number of pages, and waits between requests. It extracts page titles and H1 text for illustration; use it only where you have permission and a lawful basis. It intentionally does not attempt browser automation or access restricted pages.
Recommended Free Tools
Install the dependencies with python -m pip install requests beautifulsoup4, save the code as crawl.py, and run python crawl.py. Replace the example start URL with a site you are authorized to access and provide a real contact address in the user-agent.
from collections import deque
from urllib.parse import urljoin, urlparse, urldefrag
from urllib.robotparser import RobotFileParser
import time
import requests
from bs4 import BeautifulSoup
START_URL = "https://example.com/"
USER_AGENT = "ExampleResearchBot/1.0 (+mailto:you@example.org)"
MAX_PAGES = 10
DELAY_SECONDS = 2
start = urlparse(START_URL)
origin = f"{start.scheme}://{start.netloc}"
robots_url = origin + "/robots.txt"
robots = RobotFileParser()
robots.set_url(robots_url)
try:
robots.read()
except Exception as exc:
raise SystemExit(f"Could not read robots.txt at {robots_url}: {exc}")
session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT})
queue = deque([START_URL])
seen = set()
while queue and len(seen) < MAX_PAGES:
url = urldefrag(queue.popleft()).url
parsed = urlparse(url)
if parsed.scheme not in ("http", "https") or parsed.netloc != start.netloc:
continue
if url in seen:
continue
seen.add(url)
if not robots.can_fetch(USER_AGENT, url):
print(f"Skipped by robots.txt: {url}")
continue
try:
response = session.get(url, timeout=20)
if response.status_code in (403, 429) or response.status_code >= 500:
print(f"Stopping on HTTP {response.status_code}: {url}")
break
response.raise_for_status()
except requests.RequestException as exc:
print(f"Request failed; stopping: {url}: {exc}")
break
content_type = response.headers.get("Content-Type", "")
if "text/html" not in content_type.lower():
continue
soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else ""
heading = soup.find("h1")
h1 = heading.get_text(" ", strip=True) if heading else ""
print({"url": url, "title": title, "h1": h1})
for link in soup.select("a[href]"):
next_url = urldefrag(urljoin(url, link["href"])).url
next_parsed = urlparse(next_url)
if next_parsed.scheme in ("http", "https") and next_parsed.netloc == start.netloc:
if next_url not in seen:
queue.append(next_url)
time.sleep(DELAY_SECONDS)
This is a deliberately small example, not a complete production crawler. Python’s standard RobotFileParser is convenient for a basic demonstration, but a production system should handle robots.txt fetch outcomes and parsing in line with RFC 9309, avoid unbounded queues, and apply the site’s applicable policies. The example stops on request errors and selected HTTP failures; expand its logging, retry policy, and audit trail only in ways that remain conservative and do not turn denial responses into a reason to push harder.
What commonly goes wrong, and how should you respond?
- 403 Forbidden: The server refused the request. Stop the affected crawl and review permission, terms, and scope; do not try to evade the refusal.
- 429 Too Many Requests: The server is signaling excessive or limited traffic. Pause, reduce request rate, and respect any published guidance. If the response repeats, stop.
- Repeated 5xx responses or timeouts: The host may be unavailable or under load. Back off and stop after repeated failures; do not keep retrying at full speed.
- Pages are missing from results: Check whether the URLs were discovered, whether robots rules disallowed them, whether the response was HTML, and whether the requested fields are present in the returned content. Do not assume that adding browser automation is appropriate if access is restricted.
- Extracted fields suddenly become empty: The markup may have changed, or the content may be rendered after the initial response. Validate the parser against current pages and check for a supported API or feed before adding a browser-rendering step.
- Unexpectedly large crawl: Link graphs can lead to calendars, search results, and near-duplicate URLs. Set a page ceiling, constrain scope, normalize URLs, and provide a kill switch before running the job.
What should you budget for in a recurring crawl?
There is no defensible universal cost or volume figure for crawling: it depends on the number of permitted URLs, response size, rendering needs, schedule, and the operational controls you require. Self-hosted tooling gives you direct control over rate limits, storage, and monitoring, but your team must maintain parsers, retries, scheduling, and audit records. Managed crawling infrastructure can reduce some implementation work, but it does not remove your responsibility to confirm permission, set scope, limit traffic, and govern collected data.
For JavaScript-rendered pages or visual records, a screenshot service can be useful as a capture step, but a screenshot is not structured extraction and is not a substitute for an API or permission to access a page. ScreenshotNeo is a website screenshot API and MCP server for developers; use it when the task is to capture a clean visual page or PDF, not when you need a dataset of fields.
Best Value
Or skip the browser setup
If you have permission to capture the page and need a visual record rather than extracted fields, ScreenshotNeo takes a URL in one GET request and returns a PNG, JPEG, WebP, or PDF. Its capture flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers indicate the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for AI agents using Claude, Cursor, or another MCP client.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for setup and options. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card required.
Frequently Asked Questions
What should a crawler do if robots.txt cannot be fetched?
Do not treat every failure as permission to continue. RFC 9309 distinguishes an unavailable response from an unreachable server error; handle each deliberately, keep a record of the outcome, and use a conservative policy when the host or its instructions cannot be reached.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

