Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Use a real browser, not just an HTTP request, when images appear after JavaScript, scrolling, or interaction. The practical workflow is: open the URL in Playwright, wait for meaningful content, scroll to trigger lazy loading, inspect rendered <img> elements, collect currentSrc and responsive candidates, deduplicate them, then download each resource with safe filenames and retries. This captures images discoverable on the rendered page under your chosen viewport and interactions; it does not guarantee every asset hidden in CSS, canvas, frames, or application code.
What “all images” means in a browser
A URL can expose several different sets of images. A browser-selected set contains the resources currently chosen for the viewport and device. A declared set contains every candidate in srcset and relevant <picture><source> elements. A rendered-page scan sees images inserted by JavaScript after navigation. Static HTML parsing may miss those entirely.
Decide which result you need before writing the collector:
- What the visitor sees: collect each element’s
currentSrc. This is the URL selected by the browser, including a selectedsrcsetcandidate. - Every responsive candidate: parse
srcsetand<picture>sources as well assrc. This can produce several files for one visual. - Page-triggered attachments: listen for Playwright’s download event. That event is for downloads initiated by the page, not ordinary image resources loaded by
<img>.
Neither approach automatically discovers CSS background images, canvas pixels, images inside cross-origin frames, interaction-gated galleries, or requests made only after a particular user action. Treat the output as a documented capture of a page state, not a claim that a site contains no other images.
#1 Best Overall
Prerequisites and a safe collection plan
- Install Python 3.9 or newer and Playwright:
python -m pip install playwright, thenpython -m playwright install chromium. - Choose a destination directory and a maximum number of scroll cycles. A limit prevents a page with an unusual height or infinite feed from running forever.
- Use a reasonable delay between scrolls and downloads. Check the site’s terms, robots guidance, access controls, and image licenses before collecting or reusing files.
- Record the page URL, capture time, viewport, scroll coverage, candidate count, successful downloads, and failures. This makes the result reproducible.
Complete Playwright script in Python
The script below renders the page, scrolls in steps, re-queries the DOM, gathers the browser-selected URL plus responsive candidates, downloads each unique resource, and writes a manifest. It uses the page’s cookies and user agent when fetching images, which helps with resources that require the same session.
import asyncio
import hashlib
import json
import mimetypes
import re
from pathlib import Path
from urllib.parse import urljoin, urlparse
from playwright.async_api import async_playwright
TARGET = "https://example.com"
OUT = Path("downloaded-images")
MAX_SCROLLS = 40
SCROLL_PAUSE_MS = 700
def safe_name(resource_url, content_type=None):
parsed = urlparse(resource_url)
stem = Path(parsed.path).name or "image"
stem = re.sub(r"[^A-Za-z0-9._-]", "_", stem)[:100]
if "." not in stem:
ext = mimetypes.guess_extension((content_type or "").split(";")[0]) or ".bin"
stem += ext
digest = hashlib.sha256(resource_url.encode()).hexdigest()[:12]
return f"{stem.rsplit('.', 1)[0]}-{digest}.{stem.rsplit('.', 1)[1]}" if '.' in stem else f"{stem}-{digest}"
async def main():
OUT.mkdir(exist_ok=True)
records = {}
async with async_playwright() as pw:
browser = await pw.chromium.launch()
context = await browser.new_context(viewport={"width": 1440, "height": 1000}, device_scale_factor=1)
page = await context.new_page()
await page.goto(TARGET, wait_until="domcontentloaded", timeout=90000)
# Prefer a content-aware wait when you know the site’s selector.
await page.wait_for_timeout(1500)
previous_height = 0
for _ in range(MAX_SCROLLS):
await page.evaluate("window.scrollBy(0, Math.max(window.innerHeight * 0.8, 500))")
await page.wait_for_timeout(SCROLL_PAUSE_MS)
height = await page.evaluate("document.documentElement.scrollHeight")
if height == previous_height and await page.evaluate("window.scrollY + window.innerHeight >= document.documentElement.scrollHeight"):
break
previous_height = height
items = await page.locator("img").evaluate_all("""imgs => imgs.map(img => ({
src: img.getAttribute('src') || '',
currentSrc: img.currentSrc || '',
srcset: img.getAttribute('srcset') || '',
alt: img.getAttribute('alt') || '',
complete: img.complete,
naturalWidth: img.naturalWidth
}))""")
sources = await page.locator("picture source").evaluate_all("sources => sources.map(s => ({srcset: s.getAttribute('srcset') || ''}))")
def add(url, kind, alt="", complete=None, natural_width=None):
if not url:
return
absolute = urljoin(page.url, url)
if absolute.startswith(("data:", "blob:")):
return
records.setdefault(absolute, {"url": absolute, "kinds": set(), "alt": alt, "complete": complete, "naturalWidth": natural_width})["kinds"].add(kind)
def parse_srcset(value):
# Handles the common URL [descriptor] form; commas inside URLs are uncommon.
for part in value.split(','):
candidate = part.strip().split()[0] if part.strip() else ''
add(candidate, "srcset")
for item in items:
add(item["currentSrc"], "currentSrc", item["alt"], item["complete"], item["naturalWidth"])
add(item["src"], "src")
parse_srcset(item["srcset"])
for source in sources:
parse_srcset(source["srcset"])
cookies = await context.cookies()
cookie_header = "; ".join(f"{c['name']}={c['value']}" for c in cookies)
request = await context.request.new_context(extra_http_headers={"Cookie": cookie_header} if cookie_header else {})
manifest = []
for resource_url, info in records.items():
entry = {"url": resource_url, "alt": info["alt"], "kinds": sorted(info["kinds"]), "complete": info["complete"], "naturalWidth": info["naturalWidth"]}
try:
response = await request.get(resource_url, timeout=30000)
entry["status"] = response.status
content_type = response.headers.get("content-type", "")
entry["content_type"] = content_type
if response.ok and content_type.lower().startswith("image/"):
filename = safe_name(resource_url, content_type)
(OUT / filename).write_bytes(await response.body())
entry["file"] = filename
else:
entry["error"] = "not a successful image response"
except Exception as exc:
entry["error"] = str(exc)
manifest.append(entry)
(OUT / "manifest.json").write_text(json.dumps({"page": page.url, "images": manifest}, indent=2), encoding="utf-8")
await request.dispose()
await context.close()
await browser.close()
if __name__ == "__main__":
asyncio.run(main())
Replace TARGET, run python download_images.py, and inspect the downloaded-images directory and manifest.json. The filename includes a short hash of the source URL, so two different URLs with the same basename do not overwrite one another.
Why each discovery step matters
Wait for rendered content
domcontentloaded only means the initial document was parsed. Frameworks may insert images later. If you know a gallery or article selector, wait for that locator; otherwise use a bounded delay and repeated inspection. networkidle is not a universal completion signal because analytics, polling, and other dynamic requests can continue indefinitely.
Scroll and inspect again
Lazy-loaded images may still be pending when the page’s load event fires. Scrolling causes many sites to request the next batch. Re-query the locator after scrolling rather than retaining an early element list. For each image, complete indicates that loading finished, while naturalWidth > 0 is a useful indication that a usable image decoded; neither proves the URL is downloadable outside the browser.
Rank #2
Use currentSrc correctly
img.src can be the fallback attribute even when the browser selected a different responsive resource. currentSrc reflects that selection. If you need every offered size, parse all srcset candidates and <picture><source> values instead. These goals intentionally produce different file sets.
Resolve and filter URLs
urljoin converts relative paths to absolute URLs. Data and blob URLs are not ordinary HTTP files, so the example records neither. An extension alone is not a reliable image test; the script checks the response’s content type before writing a file.
Handling page-triggered downloads
If clicking a gallery control or export button causes an attachment download, wait for Playwright’s download event and save it explicitly:
async with page.expect_download() as download_info:
await page.get_by_role("button", name="Download").click()
download = await download_info.value
auto_name = download.suggested_filename
await download.save_as(f"downloads/{auto_name}")
Browser-context downloads are temporary and are removed when the context closes unless you save them. This event is separate from fetching the URLs used by ordinary <img> elements.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Rank #3
Images outside a basic img scan
- CSS backgrounds: inspect computed styles for
background-imageand parseurl(...)values. Styles can be injected after your first scan. - Frames: enumerate frames and scan same-origin content. Cross-origin restrictions may prevent DOM access; collect only what the site and browser policy allow.
- Canvas: a canvas has no image URL to download. If permitted, export pixels with
toDataURLortoBlob; cross-origin content can make the canvas tainted. - Interaction-gated galleries: click thumbnails, “load more,” accordions, or consent controls, then wait and scan again. Keep an action log.
- Virtualized lists: content removed as you scroll may require capturing each batch before it leaves the DOM, or observing network responses that deliver the data.
Reliability, performance, and responsible limits
Retries and failures
Record HTTP status, content type, exception text, and the source URL. Retry transient network errors and 5xx responses with a small exponential backoff, but do not retry authentication failures or persistent 4xx responses indefinitely. A browser may display an image from cache while a separate request lacks the required cookie or authorization header.
Concurrency
Downloading hundreds of files serially is slow; unrestricted concurrency can overload a site. Use a bounded semaphore, such as five simultaneous requests, and add a delay between batches. Keep the browser context alive while requests run so session cookies remain available.
Authentication and headers
For private pages, authenticate in the context before scanning. Reuse cookies and, where appropriate, the same user agent and authorization headers for resource requests. Never print tokens into the manifest or commit them to source control.
Scope reporting
Report the exact URL, viewport and device scale, interactions performed, number of scroll cycles, whether you collected currentSrc or every responsive candidate, unique URL count, successes, and failures. This is the honest boundary of “all” for a dynamic page.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #4
Common problems and fixes
The script finds too few images
Increase the bounded scroll count, wait for a known content selector, and click “load more” controls. Re-run the locator after each interaction. Check whether the page uses CSS backgrounds, frames, canvas, or a virtualized list.
Files are HTML or access-denied responses
Check the recorded content type and status. The resource may require cookies, a referer, authorization, or a signed, expiring URL. Use the authenticated browser context and do not bypass an access control that you are not authorized to access.
currentSrc is empty
The element may not have a usable source yet, may use a CSS background, or may be outside the rendered state. Scroll it into view, wait, and inspect again. For a failed load, currentSrc can still contain a URL, so check complete and naturalWidth.
The page never finishes
Remove an unbounded wait for network idle. Use a specific locator, a maximum delay, and a maximum number of scroll cycles. Dynamic sites commonly keep background requests open.
Recommended Free Tools
Best Value
Duplicate or overwritten files appear
Deduplicate normalized absolute URLs and generate collision-safe names from the full URL, as the example does. Do not trust the remote basename alone.
Or skip the browser setup
ScreenshotNeo is useful when your goal is a rendered image or PDF of a URL rather than downloading each underlying image file. It accepts one GET request and can wait, scroll through full pages, select an element, apply custom headers or cookies, and return PNG, JPEG, WebP, or PDF. Before capture it accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo documentation for all options. A direct call looks like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Sign up free for ScreenshotNeo.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesFrequently Asked Questions
Can Playwright download images that are loaded only after scrolling?
Yes. Scroll in bounded steps, wait for each batch, and re-query the page. The browser must actually reach the lazy-loaded content before its image elements and URLs can be collected.
Should I save src or currentSrc?
Use currentSrc for the resource selected for the current viewport. Parse srcset and <picture> sources when you need every responsive candidate.
Is saving an image the same as having permission to reuse it?
No. Check the image license, site terms, and applicable law for your jurisdiction and intended use; obtain authoritative legal advice for consequential cases.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

