Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scrape images from a page, request its HTML, parse every <img> element, resolve each image reference to an absolute URL, remove duplicates, then download the bytes in binary mode. The short script below handles ordinary static pages. Later sections cover lazy-loaded images, responsive srcset files, JavaScript-rendered galleries, safe filenames, retries, limits, and the legal and operational checks a real crawler needs.

What you need before collecting images

  • Python 3 and a destination directory with enough disk space.
  • requests and beautifulsoup4 for the most convenient implementation: python -m pip install requests beautifulsoup4.
  • A page you are allowed to access automatically. Check its robots.txt, terms of use, rate limits, authentication boundary, and copyright conditions first. urllib.robotparser can read crawler rules; if automated collection is disallowed, use the site’s official API or export instead.

A parser can only inspect the response it receives. It does not execute the JavaScript that may later add images to the browser DOM.

Basic one-page downloader with Requests and Beautiful Soup

This runnable example fetches a gallery, checks the response, accepts either a normal or lazy-loading source attribute, converts relative links, deduplicates them, verifies that the response is an image, and writes deterministic names.

from pathlib import Path
from urllib.parse import urljoin
import mimetypes

import requests
from bs4 import BeautifulSoup

page_url = "https://example.com/gallery"
response = requests.get(
    page_url,
    headers={"User-Agent": "image-research-bot/1.0"},
    timeout=15,
)
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")

seen = set()
out = Path("images")
out.mkdir(exist_ok=True)

for index, tag in enumerate(soup.select("img"), start=1):
    raw = tag.get("src") or tag.get("data-src")
    if not raw:
        continue

    image_url = urljoin(page_url, raw)
    if image_url in seen:
        continue
    seen.add(image_url)

    image_response = requests.get(image_url, timeout=15)
    image_response.raise_for_status()
    content_type = image_response.headers.get("content-type", "")
    if not content_type.startswith("image/"):
        continue

    extension = mimetypes.guess_extension(
        content_type.split(";", 1)[0]
    ) or ".bin"
    (out / f"image_{index:04d}{extension}").write_bytes(
        image_response.content
    )

print(f"Saved {len(seen)} unique image URLs")

response.content preserves the returned bytes, and write_bytes avoids corrupting binary files through text encoding. raise_for_status() stops on HTTP errors instead of silently saving an error page as an image.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Finding the real image URL

Normal and lazy-loaded images

Sites commonly put the first URL in src, but lazy-loading libraries may keep it in data-src, data-lazy-src, or a site-specific attribute. Extend the selection deliberately rather than downloading every attribute that happens to look like a URL.

raw = (
    tag.get("src")
    or tag.get("data-src")
    or tag.get("data-lazy-src")
    or tag.get("data-original")
)

An empty placeholder such as a transparent GIF may be the actual src while the useful file is in data-src. Inspect the HTML with your browser’s “View Source” or developer tools before choosing the attribute order.

Responsive srcset

srcset can contain several candidates, for example small.jpg 480w, large.jpg 1600w. Choose the largest width descriptor when your goal is the highest available source, then resolve it against the page URL.

from urllib.parse import urljoin

def largest_srcset_url(value, base_url):
    candidates = []
    for item in value.split(","):
        parts = item.strip().split()
        if not parts:
            continue
        width = 0
        if len(parts) > 1 and parts[1].endswith("w"):
            try:
                width = int(parts[1][:-1])
            except ValueError:
                pass
        candidates.append((width, urljoin(base_url, parts[0])))
    return max(candidates, default=(0, ""))[1]

raw = tag.get("src") or tag.get("data-src")
if tag.get("srcset"):
    raw = largest_srcset_url(tag["srcset"], page_url) or raw
if raw:
    image_url = urljoin(page_url, raw)

This obtains the largest declared candidate, not necessarily the original camera file. A thumbnail service may require a documented size parameter or an API to expose an original.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

URLs hidden in other markup

Some pages place an image in a <picture> element’s <source srcset>, Open Graph metadata, JSON-LD, or an embedded application-state object. Add selectors for those formats only when the target site documents or consistently uses them; otherwise you risk collecting unrelated preview assets.

Safer downloads for real sites

The minimal loop is intentionally small. A reusable collector should add the controls below.

Concern Practical treatment
Timeouts and transient failures Use separate connect/read timeouts, a small retry count, and exponential backoff. Log the URL and final exception.
Redirects and authentication Keep redirects enabled when appropriate, but do not cross an authentication boundary or guess credentials. Send required cookies or headers only when you are authorized.
Huge responses Stream the response and stop after a byte limit instead of loading an untrusted file into memory.
Wrong content Check the Content-Type, status, and preferably the file signature; a server can return HTML with a misleading extension.
Duplicates Normalize and store canonical URLs in a set. If query strings are tracking-only, remove them only when you know they do not identify different images.
Repeat work Persist URL, status, checksum, filename, and retrieval time in a small database or JSONL log. Add a cache and a polite delay between requests.

For streaming and a size cap:

MAX_BYTES = 20 * 1024 * 1024

with requests.get(image_url, stream=True, timeout=(5, 30)) as r:
    r.raise_for_status()
    content_type = r.headers.get("content-type", "").split(";", 1)[0]
    if not content_type.startswith("image/"):
        raise ValueError(f"Not an image: {content_type}")
    total = 0
    chunks = []
    for chunk in r.iter_content(chunk_size=64 * 1024):
        if not chunk:
            continue
        total += len(chunk)
        if total > MAX_BYTES:
            raise ValueError("Image exceeds the configured size limit")
        chunks.append(chunk)

(out / filename).write_bytes(b"".join(chunks))

For very large files, write chunks directly to a temporary file and atomically rename it after validation. Keep the original URL and response headers alongside the file so a later audit can explain how it was obtained.

Standard-library alternative with urllib

If dependencies are undesirable, urllib.request opens URLs and returns a response whose bytes can be read. Beautiful Soup can still parse those bytes, or you can use an HTML parser from the standard library.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from urllib.request import Request, urlopen
from urllib.parse import urljoin
from bs4 import BeautifulSoup

page_url = "https://example.com/gallery"
request = Request(page_url, headers={"User-Agent": "image-research-bot/1.0"})
with urlopen(request, timeout=15) as response:
    html = response.read()

soup = BeautifulSoup(html, "html.parser")
for tag in soup.select("img"):
    raw = tag.get("src")
    if raw:
        image_url = urljoin(page_url, raw)
        with urlopen(image_url, timeout=15) as image_response:
            data = image_response.read()
        # Validate headers and write data in binary mode.

Requests provides a higher-level interface, while urllib minimizes dependencies. Either client must still handle redirects, errors, limits, and access rules.

Why Beautiful Soup finds the page but not its images

The images are inserted by JavaScript

A server-rendered response may contain only a gallery shell. The browser then calls an API or runs JavaScript to insert image elements. Inspect the initial response and the Network panel. If the image data comes from an authorized JSON endpoint, call that endpoint according to its documentation. Otherwise use an authorized browser-rendering workflow; do not bypass bot checks, CAPTCHAs, or access restrictions.

The selector or attribute is wrong

Confirm that you selected img, looked for srcset and lazy-loading attributes, and resolved links with urljoin. A page can also use CSS background images, which are not img elements and require separate, site-specific parsing.

The request received a different page

Compare status, final URL, content type, and a short prefix of the response body. A consent wall, login page, redirect, or bot challenge can contain no useful images even though a browser eventually displays them.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Operational, legal, and quality checks

  • Throttle requests and honor published rate limits; a queue with bounded concurrency is safer than firing hundreds of requests at once.
  • Do not evade authentication, anti-bot controls, or explicit blocks.
  • Collecting bytes for analysis is not the same as republishing them. Check licenses, attribution requirements, privacy implications, and the terms that govern your intended use.
  • Keep a manifest containing source URL, local filename, status, content type, byte count, checksum, and timestamp.
  • Use deterministic names such as image_0001.webp or a hash-based name. Never place an unsanitized URL or title directly in a filesystem path.

Or skip the browser setup

When the page needs a rendered browser, a screenshot API can be simpler than maintaining browser automation. ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Each response identifies the page verdict and billing status with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

For a screenshot of a rendered page, make one GET request (the API can return PNG, JPEG, WebP, or PDF):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for options such as full-page capture with lazy images loaded, CSS-selector element capture, device and retina settings, custom JavaScript, waits, request blocking, headers and cookies, caching TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, PDF controls, and the usage API. Every plan includes every feature: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account to try it.

Troubleshooting checklist

  • 403 or 429: slow down, identify your client honestly, review the site’s rules, and use an official API if available.
  • 404 after parsing: resolve the reference against the document URL, not a guessed domain; preserve required query strings.
  • Files open as HTML: inspect status and Content-Type; you likely downloaded a login, error, or consent response.
  • Only thumbnails: prefer the largest srcset candidate or a documented original-image endpoint; never assume a filename rewrite is supported.
  • Memory spikes: stream with a maximum byte count and write temporary files incrementally.
  • Missing dynamically loaded images: obtain the authorized data endpoint or use a renderer; Beautiful Soup cannot execute page JavaScript.

Frequently Asked Questions

Can I scrape images from any public webpage?

No. Public visibility does not remove robots.txt, terms-of-use, rate-limit, privacy, or copyright obligations. Confirm that your intended collection and reuse are permitted.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I preserve the original file format?

Prefer the validated response Content-Type and a file-signature check over the URL suffix. Map the media type to an extension, then write the bytes unchanged.

Should I use a browser for every image-scraping task?

No. Requests or urllib plus Beautiful Soup is faster and simpler for server-rendered HTML. Use an authorized rendered-page or API approach only when JavaScript supplies the image data.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.