Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To download images from one webpage, fetch its HTML, find image references, turn relative paths into absolute URLs, then stream each response to a uniquely named file. The script below uses requests and BeautifulSoup, skips duplicates, handles collisions, checks HTTP failures, and reports responses that are not actually images.

“All images” means every image reference visible in the HTML your request receives. Images inserted later by JavaScript, protected by login, supplied through CSS backgrounds, or exposed only after a browser interaction need a different workflow.

What the downloader does

  1. Requests the page with a timeout and raises an error for an unsuccessful HTTP status.
  2. Parses the returned HTML with Beautiful Soup.
  3. Reads src values from <img> elements.
  4. Resolves root-relative, path-relative, and scheme-relative references with urllib.parse.urljoin.
  5. Deduplicates normalized URLs.
  6. Streams each image to disk in chunks instead of keeping every response in memory.
  7. Creates safe filenames and adds a suffix when two URLs would otherwise overwrite one file.
  8. Checks the response content type and records failures rather than silently saving an HTML error page as an image.

Install the Python dependencies

python -m pip install requests beautifulsoup4

The standard-library alternative is urllib.request; it requires no extra package, but Requests offers a more convenient interface for timeouts, status checks, headers, and streamed iteration.

Complete script for one permitted webpage

Save this as download_images.py. Replace the example URL with a page you are allowed to access and download from.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
  • Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.
from pathlib import Path
from urllib.parse import urljoin, urlparse, unquote
import re
import sys

import requests
from bs4 import BeautifulSoup

PAGE_URL = "https://example.com/gallery"
OUTPUT_DIR = Path("downloaded-images")
TIMEOUT = 30
CHUNK_SIZE = 1024 * 64


def safe_name(image_url: str, index: int) -> str:
    """Create a usable filename from a URL, with a deterministic fallback."""
    path_name = Path(unquote(urlparse(image_url).path)).name
    path_name = re.sub(r"[^A-Za-z0-9._-]+", "_", path_name).strip("._")
    return path_name or f"image-{index:04d}.bin"


def unique_path(directory: Path, filename: str) -> Path:
    candidate = directory / filename
    stem, suffix = candidate.stem, candidate.suffix
    number = 2
    while candidate.exists():
        candidate = directory / f"{stem}-{number}{suffix}"
        number += 1
    return candidate


def main() -> None:
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    headers = {"User-Agent": "image-downloader/1.0"}

    page_response = requests.get(PAGE_URL, headers=headers, timeout=TIMEOUT)
    page_response.raise_for_status()
    soup = BeautifulSoup(page_response.content, "html.parser")

    image_urls = []
    seen = set()
    for tag in soup.find_all("img"):
        raw_src = tag.get("src")
        if not raw_src:
            continue
        image_url = urljoin(PAGE_URL, raw_src.strip())
        if image_url not in seen:
            seen.add(image_url)
            image_urls.append(image_url)

    if not image_urls:
        print("No img[src] references were found in the returned HTML.")
        return

    print(f"Found {len(image_urls)} unique image URL(s).")
    failures = 0
    for index, image_url in enumerate(image_urls, start=1):
        try:
            with requests.get(
                image_url,
                headers=headers,
                timeout=TIMEOUT,
                stream=True,
            ) as response:
                response.raise_for_status()
                content_type = response.headers.get("Content-Type", "")
                if content_type and not content_type.lower().startswith("image/"):
                    raise ValueError(
                        f"server returned {content_type!r}, not an image"
                    )

                destination = unique_path(
                    OUTPUT_DIR, safe_name(image_url, index)
                )
                with destination.open("wb") as output:
                    for chunk in response.iter_content(CHUNK_SIZE):
                        if chunk:
                            output.write(chunk)
                print(f"[{index}/{len(image_urls)}] saved {destination}")
        except (requests.RequestException, OSError, ValueError) as exc:
            failures += 1
            print(f"[{index}/{len(image_urls)}] failed {image_url}: {exc}", file=sys.stderr)

    print(f"Finished: {len(image_urls) - failures} saved, {failures} failed.")


if __name__ == "__main__":
    main()

Run it with:

python download_images.py

Files go into downloaded-images. A query string such as ?width=1200 is not used as a filename, and duplicate basenames receive -2, -3, and later suffixes.

Why URL normalization matters

An HTML attribute may contain /media/photo.jpg, images/photo.jpg, or //cdn.example.com/photo.jpg. Concatenating strings can produce invalid addresses. urljoin(PAGE_URL, raw_src) applies the page’s scheme and directory rules correctly. Empty, missing, or malformed values are skipped or reported rather than requested blindly.

Images the basic parser does not see

srcset and lazy-loading attributes

Responsive pages may put candidates in srcset, while lazy loaders commonly use attributes such as data-src. The script intentionally starts with img[src] so its behavior is predictable. To support a known site, inspect its markup and add an explicit extraction rule for those attributes; parse each candidate and pass it through the same urljoin, deduplication, and download pipeline.

CSS backgrounds

A hero image defined in a stylesheet is not an img element. Finding it requires fetching and parsing CSS, including possible nested stylesheets, and deciding which media-query rules apply. That is site-specific and can substantially increase requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
  • Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

JavaScript-rendered content

Beautiful Soup parses the HTML returned by the server; it does not execute JavaScript. If the initial response contains an empty gallery and a script later calls an API, an HTML-only script cannot discover those images. Prefer a documented, authorized API or export when one exists. Otherwise use a browser-rendering workflow that waits for the gallery to appear, while respecting the site’s terms and access controls.

Authentication and blocked resources

Private pages may require cookies, an authorization header, or an authenticated session. Add credentials only when you are authorized to do so. A custom User-Agent can identify your client, but it does not bypass a login, CAPTCHA, rate limit, or other access control.

Handling large downloads and unreliable servers

Streaming and incomplete transfers

stream=True plus iter_content writes chunks as they arrive, limiting memory use for large files. If a connection ends early, remove the partial file or write to a temporary name and rename it only after completion. Python’s urllib.request documentation describes ContentTooShortError for an incomplete retrieval relative to a reported Content-Length; any downloader should treat an interrupted transfer as a failure, not a valid image.

Retries and rate limits

The example reports each failure and continues. For a production job, add bounded retries with backoff for transient network errors and status codes such as 429 or 503, honor any Retry-After value, and cap concurrency. A short delay between requests is safer for the host than launching hundreds at once.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
  • Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Redirects and content validation

Requests follows normal HTTP redirects. The final response can still be an HTML login page or an error document, so the content-type check is useful. It is not a cryptographic proof that bytes are a valid image: servers can mislabel content, and some legitimate image responses omit the header. For high assurance, inspect the file signature with an image library before processing it.

Standard-library version

When installing third-party packages is not possible, use urllib.request for both page and image retrieval and Beautiful Soup only if it is already available. Its urlretrieve helper can save a URL directly, but the Requests pattern gives clearer control over status checks, streaming, and cleanup. Whichever client you choose, retain URL joining, deduplication, collision-safe names, timeouts, and failure logging.

Responsible use: crawling is not permission

Google Search Central describes robots.txt this way: “A robots.txt file tells search engine crawlers which URLs the crawler can access on your site.” It is primarily a traffic and crawling-control mechanism, including for media files; it is not a security mechanism and does not decide whether you may copy or republish an image. Check the site’s terms, copyright and license information, and any applicable law before downloading or reusing files. Keep request rates reasonable.

Troubleshooting

“No img[src] references were found”

Print or save page_response.text and inspect it. You may have received a redirect, a consent/interstitial page, or a JavaScript shell. Confirm the URL and, where authorized, supply the session cookies needed for the page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
  • Easily store and access 4TB of content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

404, 403, or 429 responses

A 404 usually means the reference is stale. A 403 indicates the host rejected the request; do not treat a changed User-Agent as a guaranteed fix. A 429 means you are sending requests too quickly; slow down and follow the server’s retry guidance.

Files open as HTML

Check the logged Content-Type and final URL. The server may have returned an error or login page. The script refuses non-image content types so these responses are reported instead of hidden among your downloads.

Overwritten or missing files

Keep the collision-safe naming function, and ensure the process has write permission for the output directory. If a URL has no path filename, the indexed fallback name is used.

SSL, timeout, or connection errors

Verify the page is reachable in a browser, use a realistic timeout, and retry transient failures with backoff. Do not disable certificate verification as a routine workaround.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
UnionSine 500GB Ultra Slim Portable External Hard Drive HDD-USB 3.0
  • [Upgraded Version] - This external hard drive features a mirrored logo stripe combined with a striped anti-slip design, and the rounded corners of the casing make it easier to grip. The stripes also have a heat dissipation function, ensuring stable and fast data transfer.
  • 【Ultra-thin and quiet】 - The motherboard adopts JMicron 578 noise-free solution, giving you a quiet working environment. Lightweight and portable size designed to fit in your pocket for easy portability.
  • 【Ultra-Fast Data Transfers】 - Pairing this external hard drive with JMicron 578 solution USB 3.0 and USB 2.0 interfaces enables blazing-fast data transfer. It boasts theoretical read speeds of up to 125MB/s and write speeds of up to 103MB/s.
  • 【Plug and Play】 - With no software to install, just plug it in and the drive is ready to use.The hard disk chip is wrapped with an aluminum anti-interference layer to increase heat dissipation and protect data.
  • 【What You Get】 - 1 x Portable Hard Drive, 1 x USB 3.0 Cable, 1 x User Manual, Gift-type shell packaging ,Three-year manufacturer's warranty and free technical support services.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean capture rather than writing an HTML scraper, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—let Claude, Cursor, or another MCP client request captures.

See the ScreenshotNeo API documentation for all options. This one-call example captures the target page as WebP:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account.

Python, Node.js, and cURL calls

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Frequently Asked Questions

Can this download every image a browser displays?

No. It inventories image references present in the received HTML. JavaScript-generated images, CSS backgrounds, authenticated content, and site-specific lazy loaders require additional handling.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use Requests or urllib.request?

Requests is generally more ergonomic for timeouts, status checks, headers, and streamed writes. urllib.request avoids an extra dependency and can retrieve URLs in the standard library.

Does robots.txt give permission to reuse downloaded images?

No. It communicates crawler access preferences, not copyright permission or a license to republish.

Quick Recap

SaleBestseller No. 1
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$119.99
Bestseller No. 2
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$229.99
Bestseller No. 3
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$119.80
Bestseller No. 4
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$151.99

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.