Use Python’s requests library to fetch a page’s HTTP response, then parse its HTML with Beautiful Soup. Requests does not extract fields or run page JavaScript for you. For a reliable scraper, set a timeout, check the response status, identify your client, and make only requests the site permits.
What Requests does—and what it does not
Requests is an HTTP client: it sends requests to a server and gives your program the response. That response might contain HTML, JSON, or another resource. Requests does not interpret HTML into useful fields, and it does not execute the JavaScript a browser would run. Pair it with an HTML parser such as Beautiful Soup when the initial response contains the content you need.
In practical terms, a Requests scraper works well when the information is available in the server’s response and you can access it under the site’s rules. If a browser shows content that is absent from the response, first look for a documented API or another permitted data source. A browser-capable tool may help inspect or capture the rendered page, but a screenshot is an image, not structured text extraction.
Install the libraries and make a first request
Install Requests and Beautiful Soup in the Python environment you will use to run the script:
#1 Best Overall
python -m pip install requests beautifulsoup4
Here is a small, runnable example. It fetches a page, checks for an HTTP error, and prints its title and links. Replace the example URL with a page you are allowed to access. The CSS selectors and page structure will vary by site.
import requests
from bs4 import BeautifulSoup
url = "https://example.com/"
headers = {"User-Agent": "ExampleResearchBot/1.0 (contact: you@example.com)"}
try:
response = requests.get(url, headers=headers, timeout=(5, 20))
response.raise_for_status()
except requests.exceptions.Timeout as exc:
raise SystemExit(f"The request timed out: {exc}")
except requests.exceptions.ConnectionError as exc:
raise SystemExit(f"Could not connect: {exc}")
except requests.exceptions.HTTPError as exc:
raise SystemExit(f"The server returned an HTTP error: {exc}")
soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else "(no title)"
print("Title:", title)
for link in soup.select("a[href]"):
print(link.get_text(" ", strip=True), link["href"])
The contact value in the User-Agent is an example; use an honest identifier and a contact address you control. Do not present your client as a normal browser if it is not one. Beautiful Soup’s parser converts the HTML into a searchable tree; its selector results depend on what the server actually sent.
Build a safer scraper for repeated requests
For more than a one-off fetch, use a Session to retain cookies and reuse connections, set timeouts on every request, and handle failures explicitly. The example below fetches a list of pages, extracts a title and heading, pauses between requests, and logs failures without silently treating an error page as valid content.
import logging
import time
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
BASE_URL = "https://example.com/"
PAGES = ["/", "/about/"]
HEADERS = {
"User-Agent": "ExampleResearchBot/1.0 (contact: you@example.com)",
"Accept": "text/html,application/xhtml+xml",
}
with requests.Session() as session:
session.headers.update(HEADERS)
for path in PAGES:
url = urljoin(BASE_URL, path)
try:
response = session.get(url, timeout=(5, 20))
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else ""
heading = soup.select_one("h1")
print({
"url": response.url,
"status": response.status_code,
"title": title,
"heading": heading.get_text(" ", strip=True) if heading else "",
})
except requests.exceptions.Timeout:
logging.warning("Timeout for %s", url)
except requests.exceptions.TooManyRedirects:
logging.warning("Too many redirects for %s", url)
except requests.exceptions.HTTPError as exc:
status = exc.response.status_code if exc.response is not None else "unknown"
logging.warning("HTTP %s for %s", status, url)
except requests.exceptions.ConnectionError as exc:
logging.warning("Connection failure for %s: %s", url, exc)
except requests.exceptions.RequestException as exc:
logging.warning("Request failure for %s: %s", url, exc)
time.sleep(2)
The two-second pause is a conservative example, not a universal safe rate. Choose a rate that fits the site’s published rules and the sensitivity of its service. Keep concurrency low unless the site explicitly permits more. For a large crawl, record progress and avoid refetching unchanged pages unnecessarily.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
Why the Session matters
A Session persists cookies between related requests and can reuse underlying connections. This is useful for multi-page workflows that legitimately depend on a cookie or for repeated calls to the same host. It does not grant access the site has not authorized, and it does not turn Requests into a browser.
Parse the right response
Use response.text for decoded text such as HTML, and response.content for raw bytes such as an image or a file. Requests exposes response encoding information through response.encoding; if text appears garbled, inspect that value and the server’s response headers before overriding an encoding. For an endpoint that returns JSON, use response.json() and handle invalid JSON rather than treating every response as HTML.
Inspect a few representative pages before depending on a selector. A selector such as soup.select_one("h1") may return None when a page has no matching element or its markup changes. Check for missing values and validate the extracted output instead of assuming every page has identical structure.
Timeouts, redirects, and HTTP errors
Requests has no timeout by default. Its documentation advises using the timeout parameter in nearly all production requests; the documentation accessed in 2026 also describes release v2.34.2 and official support for Python 3.10 and later. A timeout is not a maximum total runtime for downloading a response: connect and read timeouts apply to different stages, and the elapsed wall-clock time can exceed the values you set.
Recommended Free Tools
With timeout=(5, 20), the first value limits how long to wait to establish a connection and the second limits waiting for data between reads. Tune them to your job and network, but avoid removing them. A server that accepts a connection and then stops responding can otherwise leave a script waiting indefinitely.
Requests follows redirects by default for common GET requests. The final response is available as response, and its final URL is in response.url. A redirect loop or excessive redirect chain can raise TooManyRedirects. If redirects are unexpected, inspect the response history and destination rather than blindly accepting it.
A response with an HTTP error status does not automatically become a Python exception. Call raise_for_status() when you want unsuccessful HTTP statuses to raise HTTPError. Catch the documented exception family—RequestException, including ConnectionError, Timeout, HTTPError, and TooManyRedirects—at a level where you can log the URL and failure type and choose whether to stop, skip, or retry.
Handling 403, 429, and transient failures
403 Forbidden
A 403 means the server refused the request. It may reflect access rules, authentication requirements, or automated-traffic controls. Check the site’s terms, documentation, and access requirements; confirm that you are using the intended public endpoint and appropriate credentials. Do not try to evade a block by rotating identities or disguising automation. If access is needed, ask the site owner or use an authorized API.
429 Too Many Requests
A 429 indicates that the server is rate-limiting requests. Slow down and honor a Retry-After header when present. Reduce concurrency, cache responses when freshness allows, and avoid repeatedly retrying a request that continues to be refused. A retry policy should be bounded: after a small, defined number of attempts, log the failure and stop or skip that page.
5xx responses and network failures
Server errors and intermittent connection problems can be temporary, but repeated immediate retries add load and rarely help. Use a limited retry strategy with increasing delays, and retry only failures that make sense to retry. Do not retry every 4xx response: authentication, permission, and malformed-request errors generally need a corrected request or authorized access, not another attempt.
For every retry, keep enough context to diagnose the run: requested URL, response status if available, retry count, and exception class. Avoid logging cookies, authorization headers, or other sensitive data.
Can Requests scrape a JavaScript website?
Requests cannot execute page JavaScript. It can fetch the initial HTTP response, which may include the needed content or may contain only a shell that a browser later fills in. Compare the returned HTML with the content you see in a browser. If the target data is missing from the response, inspect whether the site offers a documented, permitted API. If rendering is genuinely required, use a browser-capable approach that complies with the site’s access rules.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsBest Value
Browser automation and HTTP scraping solve different problems. Requests is lightweight for directly retrievable responses and controlled jobs; a browser has to load and render a page and may consume more time and resources. Neither approach overrides a site’s rules or guarantees access when bot checks or rate limits apply.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Respect the site’s rules and reduce unnecessary traffic
- Read the site’s terms and its
robots.txtbefore crawling. Robots directives are a useful signal about crawler access, but they are not a substitute for the site’s terms or permission where permission is required. - Identify your client honestly and provide a way to contact you where appropriate.
- Use a restrained request rate and concurrency level, and honor rate-limit signals such as 429 and
Retry-After. - Cache results when the data’s freshness requirements allow it. Avoid fetching the same page repeatedly without a reason.
- Collect only what you need, protect credentials and personal data, and stop if the site indicates that your access is not allowed.
Whether a particular scraping activity is lawful depends on the facts and applicable jurisdiction; this is not legal advice. Site terms, access controls, the type of data, and how it is used can all matter. When the stakes are material, get qualified legal guidance rather than assuming that publicly viewable means unrestricted to collect or reuse.
Or skip the browser setup
If your task is to capture a rendered page rather than extract structured text, ScreenshotNeo provides a screenshot API and MCP server. A screenshot does not replace the Requests-and-Beautiful-Soup workflow above for collecting fields, but it can be a direct option when you need an image or PDF and do not want to configure browser automation. See the ScreenshotNeo website and its API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/ -o shot.webp
Cookie and consent banners are accepted before capture, and more than 60 known consent platforms, newsletter popups, and chat widgets can be removed; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, with response headers indicating the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Sign up free for 1,000 screenshots a month, with no card required.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Troubleshooting checklist
| Symptom | Likely cause | What to check |
|---|---|---|
| The script appears to hang | No timeout, or a slow connection/read | Set a connect/read timeout such as (5, 20); log the URL and catch Timeout. |
| Output is an error page or empty result | The response was not checked, or the page structure differs | Call raise_for_status(), inspect status and final URL, then validate selectors against the returned HTML. |
| The browser shows content missing from parsed HTML | Content is added after JavaScript runs | Check for a documented permitted API or use a browser-capable method if rendering is essential. |
| 403 or CAPTCHA appears | The site denies or challenges the request | Review access rules and seek authorization; do not attempt to bypass the restriction. |
| 429 responses increase during a crawl | Request rate or concurrency is too high | Pause, honor Retry-After, lower request rate and concurrency, and use bounded retries. |
| Text is garbled | Response encoding is unexpected | Inspect response.encoding and response headers before selecting an encoding explicitly. |
| Redirect exception | Redirect loop or unusually long chain | Catch TooManyRedirects and examine the redirect destination and chain. |
What to use for each kind of page
| Need | Good starting point | Important limitation |
|---|---|---|
| HTML or JSON already in the HTTP response | Requests, with Beautiful Soup for HTML | You must check status, handle timeouts, and confirm the response contains the target data. |
| Content that appears only after JavaScript runs | A permitted API if available; otherwise a browser-capable approach | Requests alone does not run JavaScript. |
| A rendered-page image or PDF | A screenshot or browser capture tool, such as ScreenshotNeo | An image/PDF is not structured field extraction; site rules still apply. |
Frequently Asked Questions
Does Beautiful Soup send web requests?
No. It parses markup supplied to it; use an HTTP client such as Requests to fetch a response first.
What does a Requests timeout value measure?
It limits connection waiting and data-read waiting, not necessarily the total elapsed time for the complete download.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

