aiohttp can download a URL, but it does not render HTML into a PDF. A reliable Python pipeline uses aiohttp for the asynchronous HTTP request, then passes the response to a PDF renderer. Use WeasyPrint for already-rendered HTML and CSS; use Playwright when JavaScript, browser fonts, client-side data, or exact browser layout are required. Preserve an existing PDF response instead of converting it again.
What the conversion pipeline actually does
The job has two separate stages:
- Fetch: aiohttp opens the URL, follows (or restricts) redirects, checks the HTTP status, and obtains the response.
- Render: WeasyPrint or a browser engine turns HTML and CSS into paginated PDF output.
This separation matters. A page that is mostly static can be rendered directly from the downloaded HTML. A page whose content appears only after JavaScript runs needs a browser renderer. If the server already returns application/pdf, save those bytes unchanged.
Install the dependencies
Create an isolated environment and install aiohttp with the renderer you need:
python -m venv .venv
source .venv/bin/activate
python -m pip install aiohttp weasyprint
# For JavaScript-heavy pages instead:
python -m pip install aiohttp playwright
playwright install chromium
WeasyPrint also depends on native libraries that vary by operating system. Follow its installation instructions for your platform if the Python package cannot find its rendering libraries. Playwright downloads a Chromium browser when you run playwright install chromium.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11#1 Best Overall
Basic URL-to-PDF conversion with aiohttp and WeasyPrint
This reusable coroutine keeps the HTTP session scoped correctly, checks the response, records the final URL after redirects, and supplies that URL as base_url so relative stylesheets, images, and links can resolve.
import asyncio
from pathlib import Path
import aiohttp
from weasyprint import HTML
async def url_to_pdf(url: str, output: str = "out.pdf") -> None:
timeout = aiohttp.ClientTimeout(total=60)
async with aiohttp.ClientSession(timeout=timeout) as session:
async with session.get(url, allow_redirects=True) as response:
response.raise_for_status()
html = await response.text()
final_url = str(response.url)
HTML(string=html, base_url=final_url).write_pdf(output)
asyncio.run(url_to_pdf("https://example.com/"))
ClientSession is aiohttp’s recommended interface and can reuse connections through pooling and keep-alives. The nested async with blocks ensure both the session and response are closed. raise_for_status() turns a 4xx or 5xx response into an exception instead of producing a misleading PDF from an error page.
Why base_url is essential
Downloaded HTML commonly contains paths such as /styles/site.css or images/logo.svg. Once the document is represented as a string, the renderer needs a reference URL to resolve those paths. Use the final redirected URL, not necessarily the URL you originally requested.
Preserve an existing PDF
Check the response headers before rendering. If the final response is already a PDF, write the body directly:
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →async with session.get(url, allow_redirects=True) as response:
response.raise_for_status()
content_type = response.headers.get("Content-Type", "").lower()
if "application/pdf" in content_type:
with open("out.pdf", "wb") as output:
async for chunk in response.content.iter_chunked(64 * 1024):
output.write(chunk)
else:
html = await response.text()
final_url = str(response.url)
HTML(string=html, base_url=final_url).write_pdf("out.pdf")
In production, verify the detected media type and, where appropriate, inspect the first bytes for the PDF signature (%PDF-) rather than trusting a server header alone.
Stream large responses instead of loading everything
await response.text(), await response.read(), and await response.json() load the complete response into memory. For a large HTML response, stream it to a temporary file or another bounded buffer:
Rank #2
import aiohttp
async def download_html(url: str, path: str) -> None:
timeout = aiohttp.ClientTimeout(total=60)
async with aiohttp.ClientSession(timeout=timeout) as session:
async with session.get(url, allow_redirects=True) as response:
response.raise_for_status()
with open(path, "wb") as output:
async for chunk in response.content.iter_chunked(64 * 1024):
output.write(chunk)
Streaming controls memory use, but WeasyPrint still needs the document available for parsing. After downloading, you can read the temporary file and call HTML(filename=..., base_url=...), or enforce an application-level maximum before accepting a response. These size limits and timeouts are safeguards you add; they are not guarantees supplied by aiohttp.
Choose the right renderer
| Requirement | Recommended renderer | Reason |
|---|---|---|
| Server-rendered HTML and ordinary CSS | WeasyPrint | Direct HTML/CSS-to-PDF API with HTML(...).write_pdf(...). |
| JavaScript-generated content | Playwright | Runs a real browser and reproduces browser layout and client-side data loading. |
| Browser fonts, print media, interactive state | Playwright | PDF generation uses the page’s print CSS media. |
| Response is already a PDF | Neither | Save the original bytes without another conversion. |
Use Playwright for JavaScript-dependent pages
When the page fills its content with JavaScript, an aiohttp download may contain only an app shell. Playwright waits for the browser page to load and then creates the PDF:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
import asyncio
from playwright.async_api import async_playwright
async def browser_url_to_pdf(url: str, output: str = "out.pdf") -> None:
async with async_playwright() as playwright:
browser = await playwright.chromium.launch()
page = await browser.new_page()
await page.goto(url, wait_until="networkidle")
await page.pdf(path=output, print_background=True)
await browser.close()
asyncio.run(browser_url_to_pdf("https://example.com/"))
page.pdf() generates a PDF using print CSS media. networkidle is useful for pages that load data after navigation, but some sites keep analytics or live connections open indefinitely. In those cases, wait for a specific selector or use a bounded delay rather than waiting forever.
When aiohttp still belongs in a Playwright workflow
You can use aiohttp before launching the browser to perform a status check, inspect headers, supply authentication decisions, or detect that the target is already a PDF. Do not assume that a successful aiohttp request means a browser-rendered page will look identical; the two stages can receive different content when cookies, authorization, or user-agent behavior differs.
Cookies, authentication, and custom headers
Authentication must be carried deliberately from the fetch stage into rendering. The default WeasyPrint URL fetcher can open HTTP and file URLs and follows redirects, but it does not natively provide advanced cookie and authentication handling. For protected pages, either:
- Fetch the authenticated HTML and assets yourself, then provide a custom URL fetcher to WeasyPrint.
- Use a browser context in Playwright and set cookies, HTTP credentials, extra headers, or an authorization token before navigation.
- Generate an authenticated, self-contained HTML document whose resource URLs do not require a separate login.
Never place long-lived secrets in a public URL or commit them to source control. Treat URL, header, and cookie values as untrusted input and log redacted diagnostics rather than credentials.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsRedirects, limits, and safe operation
Redirect policy
The examples allow redirects because many sites move from HTTP to HTTPS or from an old path to a canonical one. If your application accepts arbitrary user URLs, restrict allowed schemes to HTTPS (and HTTP only when required), set a maximum redirect count, and validate every destination. Capture the final URL for relative-resource resolution.
Timeouts and response limits
Set an aiohttp total timeout and add application-level limits for response size, redirect count, browser navigation time, and PDF output size. A timeout should fail the job cleanly and remove any partial temporary file.
Server-side request forgery (SSRF)
A URL-to-PDF endpoint can be abused to reach private services. Apply network egress controls, reject loopback and private address ranges where appropriate, resolve and validate DNS carefully, and avoid exposing arbitrary internal responses in the generated PDF.
Improve fidelity and repeatability
- Assets: Missing images, inaccessible fonts, unsupported CSS, and cross-origin restrictions can change the result. Check that every stylesheet and image is reachable from the rendering environment.
- Print styling: Add print-specific CSS such as
@page, page margins, page breaks, andprint-color-adjustwhere supported. - Fonts: Install required fonts in the worker image or serve them from reachable URLs. Browser and WeasyPrint font substitution can differ.
- JavaScript readiness: For Playwright, wait for a meaningful selector (for example, the report container) instead of assuming navigation completion means data is ready.
- Batch jobs: Reuse one aiohttp
ClientSessionfor a batch so connections can be pooled. Bound concurrency to protect the target site and your renderer. - Reproducibility: Pin Python and renderer versions in deployment, record the final URL and status, and retain a small diagnostic copy of failed HTML when policy permits.
Troubleshooting common failures
“The PDF contains only a loading screen”
Cause: The page requires JavaScript, while aiohttp and WeasyPrint only saw the initial HTML. Fix: Switch to Playwright, wait for the data-bearing selector, and then call page.pdf().
Free tools Windows power users keep installed
One-click scans. No signup required.
“Images or CSS are missing”
Cause: Relative URLs cannot be resolved, assets require authentication, or the renderer cannot reach the host. Fix: pass the final redirected URL as base_url, make assets reachable, and provide a custom fetcher or authenticated browser context.
“ClientResponseError: 404/403/500”
Cause: The source did not return a successful page. Fix: inspect the status, redirect chain, request headers, and authentication. Do not silently render the error response.
“Timeout while waiting for network idle”
Cause: Long-lived analytics, streaming, or polling connections prevent the idle condition. Fix: wait for a specific selector or use a bounded timeout and a short post-load delay.
“WeasyPrint cannot fetch a protected resource”
Cause: Its default fetcher does not carry your application’s cookies or authorization scheme. Fix: fetch the content with aiohttp and implement a custom URL fetcher, or use Playwright with the required credentials.
“Memory usage grows during large downloads”
Cause: The whole body was read with text() or read(). Fix: consume response.content.iter_chunked(64 * 1024), enforce a size limit, and use temporary files.
Or skip the browser setup
ScreenshotNeo provides a hosted screenshot and PDF API when you do not want to operate Chromium or WeasyPrint. One GET request returns a PNG, JPEG, WebP, or PDF; its API also supports full-page capture, lazy-image loading, CSS-selector element capture, device and viewport settings, custom JavaScript and CSS, waits, cookies, headers, authorization, timezone and geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture, and a usage API. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and each response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all parameters. The following calls use the URL from the example and save the returned PDF or image bytes:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; the other listed plans are Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000). Yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to start without a card.
Best Value
FAQ
Can aiohttp itself create a PDF?
No. It handles asynchronous HTTP retrieval; a PDF renderer or an existing PDF response supplies the document layout.
Should I use WeasyPrint or Playwright for a React or Vue application?
Use Playwright when the final content depends on JavaScript execution. Use WeasyPrint only when the HTML and CSS needed for the document are already present in the fetched response.
Why does the final response URL matter after a redirect?
It gives the renderer the correct reference location for relative stylesheets, images, fonts, and links.
Recommended Free Tools
Is streaming always required?
No. Small pages are simple with await response.text(). Streaming is the safer pattern for large responses because it avoids loading the entire body at once.
Frequently Asked Questions
Can I generate a PDF from a URL without saving the HTML first?
Yes. Fetch the response text with aiohttp and pass it directly to weasyprint.HTML(string=..., base_url=...).write_pdf(...). Use a temporary file when response-size limits or memory usage make that preferable.
How can I tell whether a failed job produced a valid document?
Check the HTTP status before rendering, verify the output file exists and is non-empty, and for direct PDF responses confirm the bytes begin with the PDF signature %PDF-.
What should I test when browser and WeasyPrint PDFs differ?
Compare print CSS, installed fonts, loaded assets, JavaScript-generated content, and authentication state. They are different rendering engines, so exact visual parity is not guaranteed.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

