To extract image URLs from HTML, parse every <img> element, keep its src, expand its srcset candidates, and inspect <source> elements inside <picture>. Resolve relative references against the page URL. This produces a faithful markup inventory; it does not automatically tell you which responsive candidate a browser selected, discover every CSS background, or identify the article’s most relevant image.
Decide what “all images” means
Extraction scope changes the implementation and the result. Define it before writing code.
| Scope | What you collect | What it can miss or include |
|---|---|---|
| Markup inventory | img[src], img[srcset], and picture source references |
CSS backgrounds, JavaScript-created elements, canvas output, and resources that never appear in the fetched HTML |
| Responsive candidates | Every URL and its w or x descriptor |
It lists alternatives; it does not prove which one the current viewport uses |
| Rendered resources | Images observed by a browser after scripts, media conditions, lazy loading, and navigation run | Requires a browser and can vary with viewport, cookies, login state, and site behavior |
| Content-relevant images | A filtered set such as article figures or hero media | Relevance is a separate classification problem, not a consequence of collecting URLs |
The examples below implement a complete static-markup pass and preserve enough metadata for later browser selection. They deliberately label CSS and rendered-page coverage separately.
How HTML image references work
Single images with img
The basic case is an img element with a src attribute. Read the attribute value, then resolve a relative path such as /images/logo.png or ../hero.webp against the document URL. Keep the original value as well as the absolute URL; the original can matter for auditing or rewriting.
#1 Best Overall
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
Responsive srcset
srcset may contain several candidates separated by commas. A width descriptor looks like photo-800.jpg 800w; a pixel-density descriptor looks like photo.jpg 2x. When width descriptors are used, sizes helps the browser estimate the displayed slot. The src value can remain a fallback. Do not collapse a srcset into one arbitrary URL when the goal is a complete inventory.
picture and conditional sources
A picture element can contain multiple source elements followed by a fallback img. Each source can specify media, type, and srcset. A browser evaluates those conditions in its current environment, so a static extractor should retain the conditions and all candidates rather than claim that one URL is the currently displayed file.
CSS backgrounds and other non-markup images
Images assigned with CSS such as background-image: url(...) are outside an img-only parser. Discovering every background may require fetching linked stylesheets and examining rules or browser-computed styles. The method below reports that limitation instead of pretending to provide exhaustive CSS coverage.
Rank #2
Python extractor for img, srcset, and picture
Install the two libraries:
python -m pip install requests beautifulsoup4
Save this script as extract_images.py. It records element type, source attribute, resolved URL, descriptors, and picture conditions. It accepts an HTML document from a URL or a local file.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →from __future__ import annotations
import json
import sys
from pathlib import Path
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
def split_srcset(value: str):
"""Return [(candidate_url, descriptor_or_none)] without choosing a candidate."""
candidates = []
for part in value.split(","):
part = part.strip()
if not part:
continue
pieces = part.split()
candidate = pieces[0]
descriptor = " ".join(pieces[1:]) or None
candidates.append((candidate, descriptor))
return candidates
def extract(html: str, document_url: str):
soup = BeautifulSoup(html, "html.parser")
found = []
for img in soup.find_all("img"):
parent_picture = img.find_parent("picture")
picture_sources = []
if parent_picture:
for source in parent_picture.find_all("source", recursive=False):
raw = source.get("srcset")
if not raw:
continue
for candidate, descriptor in split_srcset(raw):
picture_sources.append({
"url": urljoin(document_url, candidate),
"original": candidate,
"descriptor": descriptor,
"media": source.get("media"),
"type": source.get("type"),
})
record = {
"tag": "img",
"alt": img.get("alt"),
"src": None,
"src_original": img.get("src"),
"srcset": [],
"sizes": img.get("sizes"),
"picture_sources": picture_sources,
}
if img.get("src"):
record["src"] = urljoin(document_url, img["src"])
if img.get("srcset"):
for candidate, descriptor in split_srcset(img["srcset"]):
record["srcset"].append({
"url": urljoin(document_url, candidate),
"original": candidate,
"descriptor": descriptor,
})
found.append(record)
return found
def load(target: str):
if target.startswith(("http://", "https://")):
response = requests.get(target, timeout=30,
headers={"User-Agent": "image-inventory/1.0"})
response.raise_for_status()
return response.text, response.url
path = Path(target).resolve()
return path.read_text(encoding="utf-8"), path.as_uri()
if __name__ == "__main__":
if len(sys.argv) != 2:
raise SystemExit("usage: python extract_images.py URL-or-file")
html, final_url = load(sys.argv[1])
print(json.dumps(extract(html, final_url), indent=2, ensure_ascii=False))
Run it with python extract_images.py https://example.com/article. A redirect is handled by using response.url as the base URL. The output keeps every candidate, its descriptor, the sizes value, and any picture conditions.
Why not select one srcset URL in Python?
Selection depends on the browser’s viewport, device pixel ratio, sizes calculation, supported format, and matching media or type. A parser can preserve the decision inputs, but a static pass cannot honestly reproduce every browser choice.
JavaScript extraction in Node.js
For server-side JavaScript, install Cheerio:
npm install cheerio
import fs from "node:fs";
import * as cheerio from "cheerio";
const html = fs.readFileSync("page.html", "utf8");
const base = "https://example.com/articles/demo";
const $ = cheerio.load(html);
const absolute = (value) => new URL(value, base).href;
const parseSrcset = (value) => value.split(",").map(x => x.trim()).filter(Boolean).map(x => {
const [url, ...descriptor] = x.split(/\s+/);
return { url: absolute(url), original: url, descriptor: descriptor.join(" ") || null };
});
const images = [];
$("img").each((_, element) => {
const img = $(element);
const picture = img.parent("picture");
const sources = [];
picture.find(":scope > source").each((_, source) => {
const s = $(source);
if (!s.attr("srcset")) return;
for (const candidate of parseSrcset(s.attr("srcset"))) {
sources.push({ ...candidate, media: s.attr("media") || null, type: s.attr("type") || null });
}
});
images.push({
src: img.attr("src") ? absolute(img.attr("src")) : null,
src_original: img.attr("src") || null,
srcset: img.attr("srcset") ? parseSrcset(img.attr("srcset")) : [],
sizes: img.attr("sizes") || null,
picture_sources: sources,
alt: img.attr("alt") || null
});
});
console.log(JSON.stringify(images, null, 2));
This has the same deliberate boundary as the Python version: it inventories references in the supplied HTML and does not execute page scripts.
Extracting from a live, JavaScript-rendered page
If the initial response contains no images but the browser later inserts them, fetch the page with a browser automation tool. Wait for a meaningful condition (for example, a selector or network idle), then inspect the DOM. Capture the final img and picture markup, and record the viewport and device scale so the result is reproducible.
- Use a real browser when images depend on client-side rendering, lazy loading, media queries, or interaction.
- Keep authentication and consent state explicit; a logged-out or consent-blocked page is a different document.
- Expect blob URLs, canvas output, extensions, and site protections to require site-specific handling. They are not guaranteed to become ordinary downloadable URLs.
For an article-image filter, collect candidates first, then apply rules such as location in the article container, dimensions, semantic attributes, or browser-rendered visibility. That is a relevance classifier, not just URL extraction.
Rank #4
Common failure modes and fixes
Relative URLs become unusable
Cause: concatenating strings or using the wrong page URL after a redirect. Fix: resolve with urljoin or the JavaScript URL constructor and use the final response URL as the base.
Only one responsive image appears
Cause: reading src and ignoring srcset or picture. Fix: preserve every candidate and its descriptor, plus each source’s media and type.
Malformed or unusual srcset
Cause: empty entries, unexpected whitespace, or authoring errors. Fix: retain the raw attribute, skip empty entries, and treat descriptors as untrusted metadata. Do not download until URLs pass your scheme and host policy.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsBest Value
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
No images in downloaded HTML
Cause: JavaScript insertion, lazy loading, authentication, or a bot/consent response. Fix: save the response for inspection, check its status and final URL, then use a browser session when the page requires rendering. Do not assume an empty result means the site has no images.
CSS images are missing
Cause: background images are not represented by img. Fix: declare CSS out of scope, or add a separate stylesheet/computed-style pass and document exactly which stylesheets and pseudo-elements you inspected.
Safety, performance, and downloading
- Prefer an allowlist of
httpandhttpsURLs. Rejectjavascript:,data:, and unexpected schemes unless your application explicitly supports them. - Limit response size, redirect count, concurrency, and per-request timeouts before downloading extracted URLs. A page can contain thousands of references.
- Deduplicate by normalized absolute URL only after retaining the original record; query strings can represent different image transformations.
- Use streaming downloads and verify the response content type when you need files, rather than trusting a filename extension.
- Cache the HTML and extraction result when repeatedly processing the same page, while respecting the site’s access rules.
Or skip the browser setup
ScreenshotNeo can render a page and return a screenshot or PDF through one request. Its cleanup step accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. It also offers an MCP server for Claude, Cursor, and other MCP clients with take_screenshot, get_page_info, and capture_pdf.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for parameters. Every plan includes its features: the Free plan provides 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots, with yearly billing providing two months free. Sign up for the free ScreenshotNeo plan when a rendered capture is more useful than setting up your own browser.
Recommended Free Tools
Choosing the right extraction method
| Need | Best starting point | Reason |
|---|---|---|
| Every URL declared in source | Python or Node parser | Fast, deterministic, and preserves markup metadata |
| Browser’s current responsive choice | Automated browser | Evaluates viewport, pixel ratio, media, type, and sizes |
| Images inserted after load | Automated browser | Sees the rendered DOM rather than only the initial response |
| CSS backgrounds | Stylesheet or computed-style inspection | Backgrounds are not img references |
| Visual proof of a page state | ScreenshotNeo | One API call handles rendering and produces a clean capture |
Frequently Asked Questions
Does extracting src download the image?
No. It only reads a reference. Downloading is a separate request that should enforce URL, size, timeout, and content-type policies.
Can an HTML parser identify the image a user sees?
Not reliably for responsive or rendered pages. It can preserve candidates and conditions; a browser is needed to evaluate the current environment.
Should I keep both relative and absolute URLs?
Yes. The absolute URL is convenient for retrieval, while the original value preserves exactly what the author placed in markup.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →

