Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteTo scrape images from a page, request its HTML, parse every <img> element, resolve each image reference to an absolute URL, remove duplicates, then download the bytes in binary mode. The short script below handles ordinary static pages. Later sections cover lazy-loaded images, responsive srcset files, JavaScript-rendered galleries, safe filenames, retries, limits, and the legal and operational checks a real crawler needs.
What you need before collecting images
- Python 3 and a destination directory with enough disk space.
requestsandbeautifulsoup4for the most convenient implementation:python -m pip install requests beautifulsoup4.- A page you are allowed to access automatically. Check its
robots.txt, terms of use, rate limits, authentication boundary, and copyright conditions first.urllib.robotparsercan read crawler rules; if automated collection is disallowed, use the site’s official API or export instead.
A parser can only inspect the response it receives. It does not execute the JavaScript that may later add images to the browser DOM.
Basic one-page downloader with Requests and Beautiful Soup
This runnable example fetches a gallery, checks the response, accepts either a normal or lazy-loading source attribute, converts relative links, deduplicates them, verifies that the response is an image, and writes deterministic names.
from pathlib import Path
from urllib.parse import urljoin
import mimetypes
import requests
from bs4 import BeautifulSoup
page_url = "https://example.com/gallery"
response = requests.get(
page_url,
headers={"User-Agent": "image-research-bot/1.0"},
timeout=15,
)
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")
seen = set()
out = Path("images")
out.mkdir(exist_ok=True)
for index, tag in enumerate(soup.select("img"), start=1):
raw = tag.get("src") or tag.get("data-src")
if not raw:
continue
image_url = urljoin(page_url, raw)
if image_url in seen:
continue
seen.add(image_url)
image_response = requests.get(image_url, timeout=15)
image_response.raise_for_status()
content_type = image_response.headers.get("content-type", "")
if not content_type.startswith("image/"):
continue
extension = mimetypes.guess_extension(
content_type.split(";", 1)[0]
) or ".bin"
(out / f"image_{index:04d}{extension}").write_bytes(
image_response.content
)
print(f"Saved {len(seen)} unique image URLs")
response.content preserves the returned bytes, and write_bytes avoids corrupting binary files through text encoding. raise_for_status() stops on HTTP errors instead of silently saving an error page as an image.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Finding the real image URL
Normal and lazy-loaded images
Sites commonly put the first URL in src, but lazy-loading libraries may keep it in data-src, data-lazy-src, or a site-specific attribute. Extend the selection deliberately rather than downloading every attribute that happens to look like a URL.
raw = (
tag.get("src")
or tag.get("data-src")
or tag.get("data-lazy-src")
or tag.get("data-original")
)
An empty placeholder such as a transparent GIF may be the actual src while the useful file is in data-src. Inspect the HTML with your browser’s “View Source” or developer tools before choosing the attribute order.
Responsive srcset
srcset can contain several candidates, for example small.jpg 480w, large.jpg 1600w. Choose the largest width descriptor when your goal is the highest available source, then resolve it against the page URL.
Rank #2
from urllib.parse import urljoin
def largest_srcset_url(value, base_url):
candidates = []
for item in value.split(","):
parts = item.strip().split()
if not parts:
continue
width = 0
if len(parts) > 1 and parts[1].endswith("w"):
try:
width = int(parts[1][:-1])
except ValueError:
pass
candidates.append((width, urljoin(base_url, parts[0])))
return max(candidates, default=(0, ""))[1]
raw = tag.get("src") or tag.get("data-src")
if tag.get("srcset"):
raw = largest_srcset_url(tag["srcset"], page_url) or raw
if raw:
image_url = urljoin(page_url, raw)
This obtains the largest declared candidate, not necessarily the original camera file. A thumbnail service may require a documented size parameter or an API to expose an original.
URLs hidden in other markup
Some pages place an image in a <picture> element’s <source srcset>, Open Graph metadata, JSON-LD, or an embedded application-state object. Add selectors for those formats only when the target site documents or consistently uses them; otherwise you risk collecting unrelated preview assets.
Safer downloads for real sites
The minimal loop is intentionally small. A reusable collector should add the controls below.
| Concern | Practical treatment |
|---|---|
| Timeouts and transient failures | Use separate connect/read timeouts, a small retry count, and exponential backoff. Log the URL and final exception. |
| Redirects and authentication | Keep redirects enabled when appropriate, but do not cross an authentication boundary or guess credentials. Send required cookies or headers only when you are authorized. |
| Huge responses | Stream the response and stop after a byte limit instead of loading an untrusted file into memory. |
| Wrong content | Check the Content-Type, status, and preferably the file signature; a server can return HTML with a misleading extension. |
| Duplicates | Normalize and store canonical URLs in a set. If query strings are tracking-only, remove them only when you know they do not identify different images. |
| Repeat work | Persist URL, status, checksum, filename, and retrieval time in a small database or JSONL log. Add a cache and a polite delay between requests. |
For streaming and a size cap:
MAX_BYTES = 20 * 1024 * 1024
with requests.get(image_url, stream=True, timeout=(5, 30)) as r:
r.raise_for_status()
content_type = r.headers.get("content-type", "").split(";", 1)[0]
if not content_type.startswith("image/"):
raise ValueError(f"Not an image: {content_type}")
total = 0
chunks = []
for chunk in r.iter_content(chunk_size=64 * 1024):
if not chunk:
continue
total += len(chunk)
if total > MAX_BYTES:
raise ValueError("Image exceeds the configured size limit")
chunks.append(chunk)
(out / filename).write_bytes(b"".join(chunks))
For very large files, write chunks directly to a temporary file and atomically rename it after validation. Keep the original URL and response headers alongside the file so a later audit can explain how it was obtained.
Standard-library alternative with urllib
If dependencies are undesirable, urllib.request opens URLs and returns a response whose bytes can be read. Beautiful Soup can still parse those bytes, or you can use an HTML parser from the standard library.
from urllib.request import Request, urlopen
from urllib.parse import urljoin
from bs4 import BeautifulSoup
page_url = "https://example.com/gallery"
request = Request(page_url, headers={"User-Agent": "image-research-bot/1.0"})
with urlopen(request, timeout=15) as response:
html = response.read()
soup = BeautifulSoup(html, "html.parser")
for tag in soup.select("img"):
raw = tag.get("src")
if raw:
image_url = urljoin(page_url, raw)
with urlopen(image_url, timeout=15) as image_response:
data = image_response.read()
# Validate headers and write data in binary mode.
Requests provides a higher-level interface, while urllib minimizes dependencies. Either client must still handle redirects, errors, limits, and access rules.
Why Beautiful Soup finds the page but not its images
The images are inserted by JavaScript
A server-rendered response may contain only a gallery shell. The browser then calls an API or runs JavaScript to insert image elements. Inspect the initial response and the Network panel. If the image data comes from an authorized JSON endpoint, call that endpoint according to its documentation. Otherwise use an authorized browser-rendering workflow; do not bypass bot checks, CAPTCHAs, or access restrictions.
The selector or attribute is wrong
Confirm that you selected img, looked for srcset and lazy-loading attributes, and resolved links with urljoin. A page can also use CSS background images, which are not img elements and require separate, site-specific parsing.
The request received a different page
Compare status, final URL, content type, and a short prefix of the response body. A consent wall, login page, redirect, or bot challenge can contain no useful images even though a browser eventually displays them.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
Operational, legal, and quality checks
- Throttle requests and honor published rate limits; a queue with bounded concurrency is safer than firing hundreds of requests at once.
- Do not evade authentication, anti-bot controls, or explicit blocks.
- Collecting bytes for analysis is not the same as republishing them. Check licenses, attribution requirements, privacy implications, and the terms that govern your intended use.
- Keep a manifest containing source URL, local filename, status, content type, byte count, checksum, and timestamp.
- Use deterministic names such as
image_0001.webpor a hash-based name. Never place an unsanitized URL or title directly in a filesystem path.
Or skip the browser setup
When the page needs a rendered browser, a screenshot API can be simpler than maintaining browser automation. ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Each response identifies the page verdict and billing status with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
For a screenshot of a rendered page, make one GET request (the API can return PNG, JPEG, WebP, or PDF):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for options such as full-page capture with lazy images loaded, CSS-selector element capture, device and retina settings, custom JavaScript, waits, request blocking, headers and cookies, caching TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, PDF controls, and the usage API. Every plan includes every feature: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account to try it.
Troubleshooting checklist
- 403 or 429: slow down, identify your client honestly, review the site’s rules, and use an official API if available.
- 404 after parsing: resolve the reference against the document URL, not a guessed domain; preserve required query strings.
- Files open as HTML: inspect status and
Content-Type; you likely downloaded a login, error, or consent response. - Only thumbnails: prefer the largest
srcsetcandidate or a documented original-image endpoint; never assume a filename rewrite is supported. - Memory spikes: stream with a maximum byte count and write temporary files incrementally.
- Missing dynamically loaded images: obtain the authorized data endpoint or use a renderer; Beautiful Soup cannot execute page JavaScript.
Frequently Asked Questions
Can I scrape images from any public webpage?
No. Public visibility does not remove robots.txt, terms-of-use, rate-limit, privacy, or copyright obligations. Confirm that your intended collection and reuse are permitted.
How do I preserve the original file format?
Prefer the validated response Content-Type and a file-signature check over the URL suffix. Map the media type to an extension, then write the bytes unchanged.
Should I use a browser for every image-scraping task?
No. Requests or urllib plus Beautiful Soup is faster and simpler for server-rendered HTML. Use an authorized rendered-page or API approach only when JavaScript supplies the image data.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

