Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To extract image URLs from HTML, parse every <img> element, keep its src, expand its srcset candidates, and inspect <source> elements inside <picture>. Resolve relative references against the page URL. This produces a faithful markup inventory; it does not automatically tell you which responsive candidate a browser selected, discover every CSS background, or identify the article’s most relevant image.

Decide what “all images” means

Extraction scope changes the implementation and the result. Define it before writing code.

Scope What you collect What it can miss or include
Markup inventory img[src], img[srcset], and picture source references CSS backgrounds, JavaScript-created elements, canvas output, and resources that never appear in the fetched HTML
Responsive candidates Every URL and its w or x descriptor It lists alternatives; it does not prove which one the current viewport uses
Rendered resources Images observed by a browser after scripts, media conditions, lazy loading, and navigation run Requires a browser and can vary with viewport, cookies, login state, and site behavior
Content-relevant images A filtered set such as article figures or hero media Relevance is a separate classification problem, not a consequence of collecting URLs

The examples below implement a complete static-markup pass and preserve enough metadata for later browser selection. They deliberately label CSS and rendered-page coverage separately.

How HTML image references work

Single images with img

The basic case is an img element with a src attribute. Read the attribute value, then resolve a relative path such as /images/logo.png or ../hero.webp against the document URL. Keep the original value as well as the absolute URL; the original can matter for auditing or rewriting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Responsive srcset

srcset may contain several candidates separated by commas. A width descriptor looks like photo-800.jpg 800w; a pixel-density descriptor looks like photo.jpg 2x. When width descriptors are used, sizes helps the browser estimate the displayed slot. The src value can remain a fallback. Do not collapse a srcset into one arbitrary URL when the goal is a complete inventory.

picture and conditional sources

A picture element can contain multiple source elements followed by a fallback img. Each source can specify media, type, and srcset. A browser evaluates those conditions in its current environment, so a static extractor should retain the conditions and all candidates rather than claim that one URL is the currently displayed file.

CSS backgrounds and other non-markup images

Images assigned with CSS such as background-image: url(...) are outside an img-only parser. Discovering every background may require fetching linked stylesheets and examining rules or browser-computed styles. The method below reports that limitation instead of pretending to provide exhaustive CSS coverage.

Python extractor for img, srcset, and picture

Install the two libraries:

python -m pip install requests beautifulsoup4

Save this script as extract_images.py. It records element type, source attribute, resolved URL, descriptors, and picture conditions. It accepts an HTML document from a URL or a local file.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from __future__ import annotations

import json
import sys
from pathlib import Path
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup


def split_srcset(value: str):
    """Return [(candidate_url, descriptor_or_none)] without choosing a candidate."""
    candidates = []
    for part in value.split(","):
        part = part.strip()
        if not part:
            continue
        pieces = part.split()
        candidate = pieces[0]
        descriptor = " ".join(pieces[1:]) or None
        candidates.append((candidate, descriptor))
    return candidates


def extract(html: str, document_url: str):
    soup = BeautifulSoup(html, "html.parser")
    found = []

    for img in soup.find_all("img"):
        parent_picture = img.find_parent("picture")
        picture_sources = []
        if parent_picture:
            for source in parent_picture.find_all("source", recursive=False):
                raw = source.get("srcset")
                if not raw:
                    continue
                for candidate, descriptor in split_srcset(raw):
                    picture_sources.append({
                        "url": urljoin(document_url, candidate),
                        "original": candidate,
                        "descriptor": descriptor,
                        "media": source.get("media"),
                        "type": source.get("type"),
                    })

        record = {
            "tag": "img",
            "alt": img.get("alt"),
            "src": None,
            "src_original": img.get("src"),
            "srcset": [],
            "sizes": img.get("sizes"),
            "picture_sources": picture_sources,
        }
        if img.get("src"):
            record["src"] = urljoin(document_url, img["src"])
        if img.get("srcset"):
            for candidate, descriptor in split_srcset(img["srcset"]):
                record["srcset"].append({
                    "url": urljoin(document_url, candidate),
                    "original": candidate,
                    "descriptor": descriptor,
                })
        found.append(record)

    return found


def load(target: str):
    if target.startswith(("http://", "https://")):
        response = requests.get(target, timeout=30,
                                headers={"User-Agent": "image-inventory/1.0"})
        response.raise_for_status()
        return response.text, response.url
    path = Path(target).resolve()
    return path.read_text(encoding="utf-8"), path.as_uri()


if __name__ == "__main__":
    if len(sys.argv) != 2:
        raise SystemExit("usage: python extract_images.py URL-or-file")
    html, final_url = load(sys.argv[1])
    print(json.dumps(extract(html, final_url), indent=2, ensure_ascii=False))

Run it with python extract_images.py https://example.com/article. A redirect is handled by using response.url as the base URL. The output keeps every candidate, its descriptor, the sizes value, and any picture conditions.

Why not select one srcset URL in Python?

Selection depends on the browser’s viewport, device pixel ratio, sizes calculation, supported format, and matching media or type. A parser can preserve the decision inputs, but a static pass cannot honestly reproduce every browser choice.

JavaScript extraction in Node.js

For server-side JavaScript, install Cheerio:

npm install cheerio
import fs from "node:fs";
import * as cheerio from "cheerio";

const html = fs.readFileSync("page.html", "utf8");
const base = "https://example.com/articles/demo";
const $ = cheerio.load(html);
const absolute = (value) => new URL(value, base).href;
const parseSrcset = (value) => value.split(",").map(x => x.trim()).filter(Boolean).map(x => {
  const [url, ...descriptor] = x.split(/\s+/);
  return { url: absolute(url), original: url, descriptor: descriptor.join(" ") || null };
});

const images = [];
$("img").each((_, element) => {
  const img = $(element);
  const picture = img.parent("picture");
  const sources = [];
  picture.find(":scope > source").each((_, source) => {
    const s = $(source);
    if (!s.attr("srcset")) return;
    for (const candidate of parseSrcset(s.attr("srcset"))) {
      sources.push({ ...candidate, media: s.attr("media") || null, type: s.attr("type") || null });
    }
  });
  images.push({
    src: img.attr("src") ? absolute(img.attr("src")) : null,
    src_original: img.attr("src") || null,
    srcset: img.attr("srcset") ? parseSrcset(img.attr("srcset")) : [],
    sizes: img.attr("sizes") || null,
    picture_sources: sources,
    alt: img.attr("alt") || null
  });
});
console.log(JSON.stringify(images, null, 2));

This has the same deliberate boundary as the Python version: it inventories references in the supplied HTML and does not execute page scripts.

Extracting from a live, JavaScript-rendered page

If the initial response contains no images but the browser later inserts them, fetch the page with a browser automation tool. Wait for a meaningful condition (for example, a selector or network idle), then inspect the DOM. Capture the final img and picture markup, and record the viewport and device scale so the result is reproducible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Use a real browser when images depend on client-side rendering, lazy loading, media queries, or interaction.
  • Keep authentication and consent state explicit; a logged-out or consent-blocked page is a different document.
  • Expect blob URLs, canvas output, extensions, and site protections to require site-specific handling. They are not guaranteed to become ordinary downloadable URLs.

For an article-image filter, collect candidates first, then apply rules such as location in the article container, dimensions, semantic attributes, or browser-rendered visibility. That is a relevance classifier, not just URL extraction.

Common failure modes and fixes

Relative URLs become unusable

Cause: concatenating strings or using the wrong page URL after a redirect. Fix: resolve with urljoin or the JavaScript URL constructor and use the final response URL as the base.

Only one responsive image appears

Cause: reading src and ignoring srcset or picture. Fix: preserve every candidate and its descriptor, plus each source’s media and type.

Malformed or unusual srcset

Cause: empty entries, unexpected whitespace, or authoring errors. Fix: retain the raw attribute, skip empty entries, and treat descriptors as untrusted metadata. Do not download until URLs pass your scheme and host policy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

No images in downloaded HTML

Cause: JavaScript insertion, lazy loading, authentication, or a bot/consent response. Fix: save the response for inspection, check its status and final URL, then use a browser session when the page requires rendering. Do not assume an empty result means the site has no images.

CSS images are missing

Cause: background images are not represented by img. Fix: declare CSS out of scope, or add a separate stylesheet/computed-style pass and document exactly which stylesheets and pseudo-elements you inspected.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Safety, performance, and downloading

  • Prefer an allowlist of http and https URLs. Reject javascript:, data:, and unexpected schemes unless your application explicitly supports them.
  • Limit response size, redirect count, concurrency, and per-request timeouts before downloading extracted URLs. A page can contain thousands of references.
  • Deduplicate by normalized absolute URL only after retaining the original record; query strings can represent different image transformations.
  • Use streaming downloads and verify the response content type when you need files, rather than trusting a filename extension.
  • Cache the HTML and extraction result when repeatedly processing the same page, while respecting the site’s access rules.

Or skip the browser setup

ScreenshotNeo can render a page and return a screenshot or PDF through one request. Its cleanup step accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. It also offers an MCP server for Claude, Cursor, and other MCP clients with take_screenshot, get_page_info, and capture_pdf.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for parameters. Every plan includes its features: the Free plan provides 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots, with yearly billing providing two months free. Sign up for the free ScreenshotNeo plan when a rendered capture is more useful than setting up your own browser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choosing the right extraction method

Need Best starting point Reason
Every URL declared in source Python or Node parser Fast, deterministic, and preserves markup metadata
Browser’s current responsive choice Automated browser Evaluates viewport, pixel ratio, media, type, and sizes
Images inserted after load Automated browser Sees the rendered DOM rather than only the initial response
CSS backgrounds Stylesheet or computed-style inspection Backgrounds are not img references
Visual proof of a page state ScreenshotNeo One API call handles rendering and produces a clean capture

Frequently Asked Questions

Does extracting src download the image?

No. It only reads a reference. Downloading is a separate request that should enforce URL, size, timeout, and content-type policies.

Can an HTML parser identify the image a user sees?

Not reliably for responsive or rendered pages. It can preserve candidates and conditions; a browser is needed to evaluate the current environment.

Should I keep both relative and absolute URLs?

Yes. The absolute URL is convenient for retrieval, while the original value preserves exactly what the author placed in markup.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.