The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →To capture every page listed in a sitemap, fetch and parse its XML, expand any sitemap index into its child files, deduplicate the page URLs, then visit each URL with one Selenium browser session and save a uniquely named PNG. The script below also records each page’s result in a CSV and continues after individual failures. Sitemap inclusion does not guarantee a URL is reachable, and Selenium’s standard screenshot captures the current browser window—not automatically the entire page.
What the script does—and what it does not
This Python example accepts a sitemap URL, handles both a <urlset> and a <sitemapindex>, and collects each <loc> address. It then opens each page in Chrome, saves a PNG, and writes a CSV manifest containing the listed URL, final URL when available, screenshot path, timestamp, and any error.
Use it only for sites you are permitted to access. A sitemap is an input list, not proof that every address works or is publicly accessible. Google describes sitemap submission as a hint; it does not guarantee that Google will fetch the sitemap or use its URLs for crawling. Google Search Central’s sitemap documentation also specifies that sitemap URLs must be absolute and documents limits of 50,000 URLs or 50 MB uncompressed per sitemap file.
Install the prerequisites
-
Install Python 3.10 or later. Selenium’s current Python documentation lists Python 3.10+ and supported browsers including Chrome, Edge, Firefox, and Safari. See Selenium’s documentation for current compatibility details.
Recommended Free Tools
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.#1 Best Overall
HP 14" HD Laptop, Windows 11, Intel Celeron Dual-Core Processor Up to 2.60GHz, 4GB RAM, 64GB SSD, Webcam, Dale Pink (Renewed)- 14" diagonal, 1366x768 resolution, HD BrightView LED, Glossy NON-TOUCH Display
-
Create and activate a virtual environment if you want to isolate the dependencies, then install Selenium and Requests:
python -m pip install selenium requests -
Make sure Chrome is installed. Selenium Manager generally handles browser-driver setup for supported platforms, so a separate driver download is often unnecessary. If setup fails in a restricted or offline environment, consult Selenium’s driver documentation and configure a compatible driver explicitly.
Run a sitemap-to-screenshots script
Save this as capture_sitemap.py. It supports ordinary XML sitemap files and sitemap indexes. It does not attempt to process compressed sitemap files or the plain-text sitemap format; decompress or convert those inputs first.
import csv
import hashlib
import re
import sys
import time
import xml.etree.ElementTree as ET
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse
import requests
from selenium import webdriver
from selenium.common.exceptions import WebDriverException
OUTPUT = Path("screenshots")
MANIFEST = OUTPUT / "manifest.csv"
REQUEST_TIMEOUT = 30
PAGE_LOAD_TIMEOUT = 60
SCRIPT_TIMEOUT = 30
def get_xml(root):
"""Download an XML sitemap and return its parsed root element."""
response = requests.get(root, timeout=REQUEST_TIMEOUT)
response.raise_for_status()
return ET.fromstring(response.content)
def local_name(tag):
"""Handle XML tags with or without a namespace."""
return tag.rsplit("}", 1)[-1]
def locs(root, expected_tag):
"""Return nonempty loc text from matching sitemap elements."""
found = []
for element in root.iter():
if local_name(element.tag) != expected_tag:
continue
for child in element:
if local_name(child.tag) == "loc" and child.text:
value = child.text.strip()
if value:
found.append(value)
return found
def collect_urls(sitemap_url, seen_sitemaps=None):
"""Expand sitemap indexes recursively and return unique page URLs."""
if seen_sitemaps is None:
seen_sitemaps = set()
if sitemap_url in seen_sitemaps:
return []
seen_sitemaps.add(sitemap_url)
root = get_xml(sitemap_url)
kind = local_name(root.tag)
if kind == "urlset":
return locs(root, "url")
if kind == "sitemapindex":
page_urls = []
for child_sitemap in locs(root, "sitemap"):
page_urls.extend(collect_urls(child_sitemap, seen_sitemaps))
return page_urls
raise ValueError(f"Expected urlset or sitemapindex, got {kind!r}")
def filename_for(index, url):
parsed = urlparse(url)
host = re.sub(r"[^A-Za-z0-9._-]+", "_", parsed.netloc) or "url"
slug = re.sub(r"[^A-Za-z0-9._-]+", "_", parsed.path.strip("/"))[:70] or "home"
digest = hashlib.sha256(url.encode("utf-8")).hexdigest()[:10]
return f"{index:05d}_{host}_{slug}_{digest}.png"
def main(sitemap_url):
OUTPUT.mkdir(parents=True, exist_ok=True)
urls = list(dict.fromkeys(collect_urls(sitemap_url)))
if not urls:
raise SystemExit("No page URLs found in the sitemap.")
options = webdriver.ChromeOptions()
# Uncomment for a headless run on a machine without a visible desktop:
# options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)
driver.set_page_load_timeout(PAGE_LOAD_TIMEOUT)
driver.set_script_timeout(SCRIPT_TIMEOUT)
fields = ["listed_url", "final_url", "screenshot", "timestamp_utc", "saved", "error"]
try:
with MANIFEST.open("w", newline="", encoding="utf-8") as csv_file:
writer = csv.DictWriter(csv_file, fieldnames=fields)
writer.writeheader()
for index, url in enumerate(urls, start=1):
path = OUTPUT / filename_for(index, url)
row = {
"listed_url": url,
"final_url": "",
"screenshot": str(path),
"timestamp_utc": datetime.now(timezone.utc).isoformat(),
"saved": False,
"error": "",
}
try:
driver.get(url)
row["final_url"] = driver.current_url
row["saved"] = driver.save_screenshot(str(path))
if not row["saved"]:
row["error"] = "save_screenshot returned False"
except (WebDriverException, Exception) as exc:
row["final_url"] = driver.current_url
row["error"] = f"{type(exc).__name__}: {exc}"
writer.writerow(row)
csv_file.flush()
print(f"{index}/{len(urls)} {url} -> {row['screenshot']} saved={row['saved']} {row['error']}")
finally:
driver.quit()
print(f"Manifest: {MANIFEST}")
if __name__ == "__main__":
if len(sys.argv) != 2:
raise SystemExit("Usage: python capture_sitemap.py https://example.com/sitemap.xml")
main(sys.argv[1])
Run it with the sitemap’s actual URL:
python capture_sitemap.py https://example.com/sitemap.xml
The script deduplicates page URLs while preserving their first-seen order. Filenames combine a sequence number, sanitized host and path, and a short hash of the full URL, reducing collisions when paths are similar. The CSV is flushed after every row so results already recorded remain available if the run is interrupted.
Wait for dynamic pages and choose screenshot dimensions
Wait for content that appears after navigation
driver.get() returning does not establish that every site-specific widget, client-rendered section, or lazy-loaded image is ready. Add a condition that reflects the target site, such as waiting for a known content selector:
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
# After driver.get(url):
WebDriverWait(driver, 20).until(
lambda d: d.find_element(By.CSS_SELECTOR, "main article").is_displayed()
)
Replace the selector with one that exists on the pages being captured, and handle a missing selector as a per-URL failure or a recorded readiness timeout. A fixed sleep can be used as a deliberate delay, but it is not a guarantee that a page has finished rendering.
Rank #2
- 1.1 GHz (boost up to 2.4GHz) Intel Celeron N5030 Quad-Core
- 4GB DDR4 System Memory; 128GB Solid State Drive
- 11.6" HD (1366 x 768) Multi-Touch Display
- Combo headphone/microphone jack - Noble Wedge Lock slot - HDMI; 2 USB 3.1 Gen 1
- Windows 11 Pro
Set viewport and understand the capture area
Set a consistent viewport before the loop if screenshots must be comparable:
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minutedriver.set_window_size(1365, 900)
Selenium’s save_screenshot(path) saves the current window as a PNG and returns a boolean indicating whether the save succeeded. It is not, by itself, a full-page screenshot operation. The capture behavior is described in Selenium’s windows and tabs documentation and the Python WebDriver API reference. For full-page output, use a separately supported browser-specific technique or another capture method, and verify the resulting dimensions rather than assuming the standard call includes content below the viewport.
Choose whether to reuse or restart the browser
The example reuses one Chrome session, which avoids repeated startup overhead and is a practical default for a straightforward batch. A reused session also carries cookies, local storage, and other page state from one visit to the next. If that state affects results, isolate captures or restart the driver periodically; doing so costs additional startup time and resources. These are workflow trade-offs, not guarantees about Selenium behavior.
For visual comparisons over time, record or pin the Selenium and browser versions, viewport, locale, timezone, and authentication state. Otherwise, a difference in environment can look like a change to the page itself.
Scale the run without overwhelming the site or machine
Google’s documented per-file sitemap limits are 50,000 URLs or 50 MB uncompressed, but a browser visit for every page can take substantial time and disk space. Estimate the output size and runtime on a small sample before running a large list. Use a reasonable request pace and comply with the site’s access rules; a sitemap does not guarantee that each URL is reachable or that bot checks will allow access.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute-
Keep the default sequential loop for a modest batch. It uses one browser and avoids launching an unbounded number of browser processes.
-
If you need concurrency, design an explicit worker and resource budget. Each independent browser consumes memory and CPU, and concurrent requests can place more load on the target site.
Rank #3
Dell Latitude 5420 14" FHD Business Laptop Computer, Intel Quad-Core i5-1145G7, 16GB DDR4 RAM, 256GB SSD, Camera, HDMI, Windows 11 Pro (Renewed)- 256 GB SSD of storage.
- Multitasking is easy with 16GB of RAM
- Equipped with a blazing fast Core i5 2.00 GHz processor.
-
Retain the manifest. It provides a map from each listed address to its output file and makes failed pages easy to retry without recapturing successful ones.
-
For recurring captures, decide whether a redirect should be identified by its original sitemap URL or final destination. The script records both, while keeping the original URL in the filename input and manifest.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Troubleshooting
The script cannot download or parse the sitemap
-
HTTP error or timeout: check the sitemap address, network access, authentication requirements, and server response. Increase
REQUEST_TIMEOUTonly when a slow response is expected. -
XML parse error: confirm that the response is actually XML rather than an HTML error page. If the file is compressed, decompress it before parsing; this example does not handle compressed input.
-
Unexpected root element: the code accepts XML roots named
urlsetandsitemapindex. A plain-text list or a different XML format needs its own parser. -
No URLs found: inspect the sitemap contents and confirm that each page entry has a nonempty
locelement. The code does not discover a sitemap fromrobots.txt; supply the sitemap URL directly.Free tools Windows power users keep installed
One-click scans. No signup required.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Chrome or Selenium will not start
-
Install a supported Chrome browser and confirm that the environment permits Selenium Manager to obtain or locate the matching driver.
Rank #4
15.6 Inch Laptop Computer, N4020, 4GB DDR4 RAM, 128GB eMMC,with Windows 11- EFFORTLESS EVERYDAY PERFORMANCE: Powered by Intel Celeron N4020 processor and Windows 11 Home system, delivering reliable, low-power efficiency for daily tasks like document editing, email, online classes, and web browsing
- 15.6-INCH FULL HD DISPLAY: Enjoy immersive visuals on the 15.6" FHD (1920x1080) anti-glare screen with micro-edge bezels. Delivers clear details and comfortable viewing for long study sessions, working on spreadsheets, and video playback
- RESPONSIVE MULTITASKING & STORAGE: Built with 4GB LPDDR4 RAM and 128GB eMMC storage for smooth daily essential use. Expand your storage by up to 1TB via the integrated TF card slot to easily store movies, photos, and working files
- ADVANCED CONNECTIVITY: Outfitted with 2x Full-Featured Type-C ports for data transfer, fast charging, and dual-monitor output, alongside 2x USB 3.2 Gen1 ports and a 3.5mm audio jack for complete peripheral compatibility
- LIGHTWEIGHT & SILENT OPERATION: Slim and portable for effortless travel or commuting. Features a 1MP HD webcam for remote meetings, 38Wh battery with 45W Type-C fast charging, and a fanless silent design for peaceful work environments.
-
In a headless or server environment, uncomment the headless option. If browser startup still fails, check the browser’s own startup error and Selenium’s current platform guidance.
A page times out, redirects, or shows a challenge
-
A page-load timeout is recorded for that URL and the loop continues. Check the manifest’s error and final URL, then determine whether the failure is temporary, access-controlled, or caused by site behavior.
-
A redirect is not automatically an error: the manifest preserves both the sitemap URL and the browser’s final URL.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy. -
Authentication, geoblocking, robots rules, network failures, and bot protections may prevent access. Only configure authentication or other browser state when authorized to do so.
The screenshot is blank, incomplete, or the file is missing
-
Wait for a page-specific readiness condition if content renders after navigation; the page-load event alone may be insufficient.
-
Check
savedanderrorin the manifest. Selenium’s screenshot API returns a boolean; a false result is recorded rather than treated as a successful image. -
Confirm the output directory is writable and that available disk space is sufficient. Standard WebDriver screenshots capture the current window, so content outside the viewport may not appear.
Recommended: PC Feels Slow? A Free Scan Shows What's Dragging Windows Down →Recommended: Update Every Outdated Driver on Your PC in One Scan - Free →Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.Best Value
Sale15.6 Inch Win 11 Laptop Computer, N4020, 4GB DDR4 RAM, 128GB Storage- WINDOWS 11 | STABLE PERFORMANCE: Powered by Intel Celeron N4020 processor and Windows 11 system, this laptop delivers stable performance for everyday computing tasks. It supports web browsing, online learning, document editing, email communication, and basic office work with optimized power efficiency, providing a practical and reliable experience for essential daily use for daily use.
- 15.6” FHD IPS DISPLAY: Features a 15.6-inch Full HD IPS display with narrow bezels, offering wider viewing angles and clearer image details compared to standard panels. The improved screen-to-body ratio enhances visual experience for study, reading, document work, and video playback, making it suitable for both productivity and entertainment use.
- 4GB DDR4 + 128GB eMMC STORAGE: Equipped with 4GB DDR4 memory and 128GB eMMC storage for everyday basics such as browsing, documents, email, and online learning platforms. The built-in TF card slot supports storage expansion up to 1TB, giving you more flexibility for files, photos, videos, and daily documents. TF card not included.
- CONNECTIVITY & PORTS: Includes 1× TF card slot, 2× USB 3.2 Gen1 ports, and 2× full-featured Type-C ports (USB 3.2 Gen1). The Type-C ports support data transfer, charging, and video output, enabling flexible connection with external devices such as monitors, storage, and peripherals for daily work and study use.
- LIGHTWEIGHT DESIGN | ONLINE COMMUNICATION: Designed with a slim, portable profile, this laptop is easy to carry for school, commuting, and travel. A built-in 1MP front camera supports online classes, video meetings, remote communication, and everyday conferencing. The 3300mAh battery works with the low-power system design to support practical daily use, while thermal optimization helps maintain quieter operation during extended tasks.
Or skip the browser setup
ScreenshotNeo offers a one-call screenshot API if you do not want to manage Selenium and a browser locally. For a single URL, for example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for request options and response details. For sitemap-scale work, your code still needs to parse the sitemap and make a request for each page; the call above captures one URL.
-
Cookie banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
-
Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy. -
An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents and MCP clients. -
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
Frequently Asked Questions
Does this script capture every URL in a sitemap index?
Yes. It recursively expands sitemap indexes and deduplicates the page URLs it collects from their child sitemap files.
Does Selenium’s save_screenshot capture a full page?
No. The standard WebDriver screenshot saves the current browser window. Use and verify a separate full-page capture technique if you need content beyond the viewport.
Can I use this with a compressed sitemap?
Not directly. Decompress the sitemap first; the example handles XML urlsets and indexes, not compressed files or plain-text URL lists.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

