iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Use Python’s csv module to read the URLs, Playwright to capture each page, and Pillow to resize and arrange the images into a labeled grid. The script below keeps CSV order, records each result in a manifest, and continues when an individual URL fails. It creates viewport screenshots by default; switch one setting for full-page captures.
What you need
- Python 3 and a CSV file with a header named
url. - Playwright and its Chromium browser for page navigation and screenshots.
- Pillow to build the contact sheet.
The example assumes a comma-delimited UTF-8 CSV. If the file uses another delimiter or encoding, adjust the configuration noted below. It writes screenshots to screenshots/, a traceable manifest.csv, and a composite contact_sheet.png.
Install the dependencies
-
Create and activate a virtual environment if desired, then install the Python packages:
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.python -m pip install playwright pillow python -m playwright install chromium -
Save the following script as
capture_csv.pybeside your input file, or editINPUT_CSVto point to it.#1 Best Overall
SaleEpson Workforce ES-50 Compact & Lightweight Mobile Document Scanner- PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
- QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
- VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
- INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
- EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
Run the complete capture and contact-sheet script
import csv
import re
from pathlib import Path
from urllib.parse import urlparse
from PIL import Image, ImageDraw, ImageFont
from playwright.sync_api import sync_playwright
INPUT_CSV = Path("urls.csv")
URL_COLUMN = "url"
OUTPUT_DIR = Path("screenshots")
MANIFEST = Path("manifest.csv")
CONTACT_SHEET = Path("contact_sheet.png")
# Viewport screenshots make more comparable tiles. Set True to capture
# the entire scrollable page instead.
FULL_PAGE = False
# Navigation waits until the page's load event, up to this many milliseconds.
NAVIGATION_TIMEOUT_MS = 30_000
# Contact-sheet layout, in pixels.
THUMB_WIDTH = 320
THUMB_HEIGHT = 200
LABEL_HEIGHT = 48
COLUMNS = 3
PADDING = 16
def safe_filename(index, url):
host = urlparse(url).netloc or "page"
host = re.sub(r"[^A-Za-z0-9.-]+", "-", host).strip("-.") or "page"
return f"{index:04d}-{host}.png"
def make_contact_sheet(items):
if not items:
print("No successful screenshots; contact sheet not created.")
return
rows = (len(items) + COLUMNS - 1) // COLUMNS
tile_width = THUMB_WIDTH
tile_height = THUMB_HEIGHT + LABEL_HEIGHT
sheet = Image.new(
"RGB",
(PADDING + COLUMNS * (tile_width + PADDING),
PADDING + rows * (tile_height + PADDING)),
"white",
)
draw = ImageDraw.Draw(sheet)
font = ImageFont.load_default()
for position, item in enumerate(items):
row, col = divmod(position, COLUMNS)
x = PADDING + col * (tile_width + PADDING)
y = PADDING + row * (tile_height + PADDING)
with Image.open(item["path"]) as source:
thumb = source.convert("RGB")
thumb.thumbnail((THUMB_WIDTH, THUMB_HEIGHT))
# Center the aspect-preserving thumbnail in a fixed preview box.
left = x + (THUMB_WIDTH - thumb.width) // 2
top = y + (THUMB_HEIGHT - thumb.height) // 2
sheet.paste(thumb, (left, top))
label = f"{item['index']}: {item['url']}"
draw.text((x, y + THUMB_HEIGHT + 4), label[:72], fill="black", font=font)
draw.text((x, y + THUMB_HEIGHT + 20), item["status"], fill="black", font=font)
sheet.save(CONTACT_SHEET, format="PNG")
def main():
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
records = []
# newline='' is recommended for CSV files. utf-8-sig also accepts files
# that begin with a UTF-8 byte-order mark.
with INPUT_CSV.open("r", newline="", encoding="utf-8-sig") as csv_file:
reader = csv.DictReader(csv_file)
if not reader.fieldnames or URL_COLUMN not in reader.fieldnames:
raise ValueError(
f"CSV must have a header named {URL_COLUMN!r}; "
f"found {reader.fieldnames!r}"
)
with sync_playwright() as playwright:
browser = playwright.chromium.launch()
page = browser.new_page(viewport={"width": 1365, "height": 900})
page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)
for index, row in enumerate(reader, start=1):
original_url = row.get(URL_COLUMN) or ""
url = original_url.strip()
record = {
"index": index,
"url": original_url,
"filename": "",
"status": "",
"error": "",
}
if not url:
record["status"] = "skipped: blank URL"
records.append(record)
print(f"Row {index}: blank URL; skipped")
continue
filename = safe_filename(index, url)
path = OUTPUT_DIR / filename
record["filename"] = str(path)
try:
response = page.goto(url, wait_until="load")
page.screenshot(path=str(path), full_page=FULL_PAGE)
if response is not None and response.status >= 400:
record["status"] = f"captured: HTTP {response.status}"
else:
record["status"] = "captured"
print(f"Row {index}: {record['status']} - {url}")
except Exception as exc:
record["status"] = "failed"
record["error"] = str(exc).replace("n", " ")
print(f"Row {index}: failed - {url} - {record['error']}")
records.append(record)
browser.close()
with MANIFEST.open("w", newline="", encoding="utf-8") as manifest_file:
writer = csv.DictWriter(
manifest_file,
fieldnames=["index", "url", "filename", "status", "error"],
)
writer.writeheader()
writer.writerows(records)
successful = [
{"index": r["index"], "url": r["url"].strip(),
"path": r["filename"], "status": r["status"]}
for r in records if r["status"].startswith("captured")
]
make_contact_sheet(successful)
print(f"Manifest: {MANIFEST}")
print(f"Successful captures: {len(successful)} of {len(records)} CSV rows")
if __name__ == "__main__":
main()
Run it with:
python capture_csv.py
The index follows data-row order, beginning at 1 after the header. Each screenshot filename starts with that index, so duplicate hosts do not overwrite one another. The manifest preserves the original URL cell, output filename, status, and error text. Failed and blank rows remain in the manifest; only successful captures appear as contact-sheet tiles.
Choose what each screenshot shows
Viewport or full page
With FULL_PAGE = False, each image shows the same browser viewport, making page previews easier to compare in a grid. Set it to True to include the full scrollable page. Playwright describes a full-page screenshot as one where the page is captured “as if you had a very tall screen and the page could fit it entirely” in its Python screenshots guide. Long pages shrink substantially when fitted into a thumbnail, so a viewport is often more useful for scanning while full-page images are better when below-the-fold content matters.
Rank #2
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
Wait behavior
This script uses wait_until="load", then captures immediately after the load event. Sites with delayed content, animations, or client-rendered data may need a different wait strategy. Playwright’s navigation API documents the navigation options; use an explicit selector wait or a short delay only when the target pages need it. A longer timeout may help genuinely slow pages but makes a batch take longer when pages hang.
Image format and thumbnails
Playwright supports PNG, JPEG, and WebP screenshots through its screenshot API. This workflow saves PNGs and a PNG contact sheet, a sensible default for text-heavy previews. JPEG or WebP can reduce image size where storage matters, though the contact sheet remains a separate composite. The thumbnail dimensions, three-column layout, and label clipping in this example are adjustable constants, not universal best settings.
Rank #3
- STAY ORGANIZED – Easily convert your paper documents into digital formats like searchable PDF files, JPEGs, and more.Power Consumption : 2.5W or less (Energy Saving Mode: 0.7W). Suggested Daily Volume : 500 scans..Does it contain liquid: no
- CONVENIENT AND PORTABLE –lightweight and small in size, you can take the scanner anywhere from home offices, classrooms, remote offices, and anywhere in between
- HANDLES VARIOUS MEDIA TYPES – Digitize receipts, business cards, plastic or embossed cards, reports, legal documents, and more
- FAST AND EFFICIENT – No technical hurdles or complicated setups here; easily scan both sides of a document at the same time, in color or black-and-white, at up to 12 pages-per-minute, and with a 20 sheet automatic feeder
- BROAD COMPATIBILITY – Works with both Windows and Mac devices, be it laptop or computer
Adapt the CSV and batch behavior
Use another URL header
Change URL_COLUMN = "url" to the exact column heading in your file, such as "Website". Header names are matched exactly, including capitalization and spaces.
Use a different delimiter or encoding
Python’s CSV documentation notes that CSV files can differ in formatting across applications. For a semicolon-separated file, pass delimiter=";" when creating csv.DictReader. For an encoding other than UTF-8, change encoding="utf-8-sig" to the appropriate encoding for the file. Keep newline="" when opening it.
Rank #4
- IRIScan Express, portable scanner : scans color and black and white documents a blazing speed up to 8ppm simplex. Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- IRIScan Express mobile scanner is powered via an included micro USB 2. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan. USB cable provided. AC Adapter not provided and not needed.
- IRIScan flatbed scanner uses a simplex scanning mode allows for quick and straightforward scanning of single-sided documents. IRIScan with its full portable features is the ideal document scanners for computers.
- IRIScan document scanner : Versatile scanning capabilities, including scanning to Word, PDF, and Excel formats with companion software provided Readiris OCR
- Receipt scanner and card scanner with Additional features include scanning business cards directly to Outlook, photo scanning, and receipt scanning for efficient document management
Keep the sheet readable at scale
A single image gets unwieldy as the URL count grows. Increase COLUMNS only if labels and previews remain readable; otherwise reduce the number of successful items per sheet and create multiple sheets, or use a paginated HTML gallery. For very long URLs, replace the label with a shortened host/path or row number and use manifest.csv to look up the full address.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Troubleshoot common failures
- Missing URL header: the script stops before navigating if there is no exact
urlheader. Rename the CSV header or changeURL_COLUMN. - Blank or malformed rows: blank values are skipped and recorded. For a malformed address, confirm it includes a scheme such as
https://; some sites redirect or reject automation regardless. - Navigation timeout: the row is marked failed and later rows continue. Check the URL and network access, then increase
NAVIGATION_TIMEOUT_MSif the site is merely slow. Do not treat a longer timeout as a fix for a page that never completes. - HTTP error response: a page may return a screenshot despite an HTTP status of 400 or higher. The manifest labels it
captured: HTTP ...; inspect the page and status rather than assuming the image represents a successful page. - Browser executable missing: run
python -m playwright install chromiumin the same Python environment where Playwright is installed. - Contact sheet omits a row: the sheet includes captures whose status starts with
captured; failed and blank rows are intentionally not converted into image tiles. Review their manifest entries. - Tiny or unreadable tile content: use viewport captures for consistent previews, enlarge
THUMB_WIDTHandTHUMB_HEIGHT, or split the output into more sheets.
Performance, reliability, and cost considerations
The script processes one URL at a time in a single browser page. That keeps the mapping between input rows and results straightforward, but total run time depends on site response, chosen waits, timeouts, and page length; there is no meaningful universal speed figure. Serial processing also avoids sending a burst of concurrent visits to every host. If you later add concurrency, keep per-row error handling and consider each target site’s access policies and capacity.
Best Value
- Scanner type: Document
- Connectivity technology: USB
- With Auto Scan Mode, the scanner automatically detects what you're scanning
- Digitize documents and images
Browser automation cannot guarantee that every address will permit access or render the same way as a human visit. Redirects, bot checks, consent prompts, authentication requirements, network failures, and dynamic content can affect the result. The manifest makes these outcomes visible instead of silently dropping rows. Capturing locally has no per-screenshot service charge, but it uses your machine’s CPU, memory, network, and storage.
Or skip the browser setup
For one-request-per-URL capture through an API, ScreenshotNeo accepts a URL and returns an image or PDF. Its API can also make clean shots by accepting consent banners and removing more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server provides screenshot tools for Claude, Cursor, and other MCP clients.
Example cURL request (replace the URL with a value from your CSV):
Free tools Windows power users keep installed
One-click scans. No signup required.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for authentication, output options, and request parameters. ScreenshotNeo has a free plan with 1,000 shots per month and no card; paid plans start at $5 for 3,000 shots. Sign up at ScreenshotNeo to start with the free allowance.
Frequently asked questions
Does the contact sheet include failed URLs?
No. Failures and blank entries are recorded in the manifest, while the contact sheet contains successful screenshot files only.
Can I use the output images in another format?
Yes. Playwright can save screenshots as PNG, JPEG, or WebP; choose the format in the screenshot call and adjust the filename extension to match.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

