Use aiohttp to download the PDF and pypdf to select and write pages. For a small file you can read the response into memory; for a large file, stream response.content to disk in chunks, then add validated, zero-based page indexes to a new PdfWriter.
What each library does
This workflow has two separate jobs:
- aiohttp performs the asynchronous HTTP request and transfers the bytes.
- pypdf parses the PDF, exposes its pages, and creates the selected-pages document.
aiohttp is not a PDF editor, and pypdf does not download URLs. Keeping those responsibilities separate makes failures easier to diagnose.
Install the dependencies
python -m pip install aiohttp pypdf
The examples use the current APIs represented by aiohttp 3.14.3 documentation and pypdf 6.4.2 documentation. Check the documentation for the versions pinned in your project if you are adapting the code.
Complete example: download and export selected pages
This script downloads a PDF to a temporary local path, checks the HTTP response, validates the requested indexes, and writes a second PDF.
Recommended Free Tools
#1 Best Overall
import asyncio
from pathlib import Path
from tempfile import TemporaryDirectory
import aiohttp
from pypdf import PdfReader, PdfWriter
CHUNK_SIZE = 64 * 1024
async def download_pdf(
session: aiohttp.ClientSession,
url: str,
destination: Path,
) -> None:
async with session.get(url) as response:
response.raise_for_status()
with destination.open("wb") as output:
async for chunk in response.content.iter_chunked(CHUNK_SIZE):
output.write(chunk)
def export_pages(
source: Path,
destination: Path,
page_indexes: list[int],
) -> None:
reader = PdfReader(source)
page_count = len(reader.pages)
invalid = [
index for index in page_indexes
if index < 0 or index >= page_count
]
if invalid:
raise ValueError(
f"Invalid page indexes {invalid}; PDF has {page_count} pages"
)
writer = PdfWriter()
for index in page_indexes:
writer.add_page(reader.pages[index])
with destination.open("wb") as output:
writer.write(output)
async def main() -> None:
url = "https://example.com/document.pdf"
selected_pages = [0, 2, 3] # human pages 1, 3, and 4
with TemporaryDirectory() as temporary_directory:
source = Path(temporary_directory) / "input.pdf"
await download_pdf(url, source)
export_pages(source, Path("selected-pages.pdf"), selected_pages)
if __name__ == "__main__":
asyncio.run(main())
Replace the example URL with the actual PDF URL. The resulting selected-pages.pdf contains the first, third, and fourth human-numbered pages.
Convert human page numbers to Python indexes
People normally count a PDF from page 1; Python sequences start at index 0. Convert at the boundary of your application:
- Human page 1 becomes index 0.
- Human pages 1, 3, and 4 become indexes
[0, 2, 3]. - Human pages 2 through 5 become indexes
1, 2, 3, 4.
For a contiguous human-inclusive range, the equivalent half-open Python slice is start - 1:end. When adding pages individually, use range(start - 1, end):
human_start, human_end = 2, 5
indexes = range(human_start - 1, human_end)
for index in indexes:
writer.add_page(reader.pages[index])
This is Python indexing logic, not a special range syntax supplied by pypdf. Validate both endpoints before indexing.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Small-download alternative: read the response into memory
For a modest PDF, an in-memory approach is shorter. aiohttp’s read(), json(), and text() convenience methods load the whole response into memory, so this version trades simplicity for memory consumption.
import asyncio
from io import BytesIO
import aiohttp
from pypdf import PdfReader, PdfWriter
async def main() -> None:
async with aiohttp.ClientSession() as session:
async with session.get("https://example.com/document.pdf") as response:
response.raise_for_status()
data = await response.read()
reader = PdfReader(BytesIO(data))
writer = PdfWriter()
for index in (0, 2, 3):
writer.add_page(reader.pages[index])
with open("selected-pages.pdf", "wb") as output:
writer.write(output)
asyncio.run(main())
Do not assume that streaming makes total memory usage constant: pypdf still has to parse the downloaded document, and PDF structure can require substantial memory.
Rank #2
Stream safely for large files
The chunked example avoids constructing one giant Python bytes object for the HTTP body. The async with blocks close the response and session, while the file context manager flushes and closes the destination.
Set practical request limits
Production code should define a timeout and, when the application accepts arbitrary URLs, enforce a maximum download size. A timeout prevents a stalled server from occupying a task indefinitely:
timeout = aiohttp.ClientTimeout(total=90)
async with aiohttp.ClientSession(timeout=timeout) as session:
...
A size limit is an application policy rather than an aiohttp guarantee. Track bytes while iterating and abort when the limit is exceeded.
MAX_BYTES = 200 * 1024 * 1024
received = 0
async for chunk in response.content.iter_chunked(64 * 1024):
received += len(chunk)
if received > MAX_BYTES:
raise ValueError("PDF exceeds the configured size limit")
output.write(chunk)
Validate the response before treating it as a PDF
response.raise_for_status() converts 4xx and 5xx responses into exceptions instead of silently saving an error page. A successful status still does not prove the body is a valid PDF: servers can return HTML with status 200, redirect to a login page, or send a truncated file.
- Inspect the final response URL when redirects matter.
- Optionally check that the
Content-Typeis compatible with a PDF, while allowing servers that omit or mislabel it. - After download, let
PdfReaderparse the file and report malformed input rather than publishing the output as valid.
Handle duplicate pages, ordering, and empty selections
PdfWriter.add_page follows the order in which you call it. You can export pages in any order and can deliberately repeat an index:
for index in (4, 1, 1, 7):
writer.add_page(reader.pages[index])
An empty selection should be rejected by your application, because writing a zero-page PDF is rarely what a user intended:
if not page_indexes:
raise ValueError("Select at least one page")
Common errors and fixes
401, 403, or 404 from aiohttp
The URL may require authentication, block your client, or be wrong. Confirm the URL, supply the required headers or cookies, and inspect the server’s response. Keep raise_for_status() enabled so the failure is visible.
ContentTypeError or an HTML file saved as a PDF
This often means the URL returned a login page, consent page, or error document. Follow redirects deliberately, authenticate if required, and inspect the saved bytes before passing them to pypdf.
IndexError while accessing reader.pages[index]
The requested index is negative or at least the page count. Read len(reader.pages), convert human numbers correctly, and perform explicit validation before adding pages.
pypdf cannot read the file
The input may be encrypted, malformed, truncated, or unusually complex. Preserve the original download for diagnosis, verify that the transfer completed, and handle encryption according to the document owner’s permissions. Do not promise that every PDF can be parsed without additional handling.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →The process uses too much memory
Switch from await response.read() to chunked disk writing, reduce concurrency, and impose a file-size limit. Remember that parsing and writing the PDF can still consume memory.
The output is unexpectedly large
pypdf copies page objects and their referenced resources. A selected page can retain fonts, images, and other objects from the source. Avoid assuming that selecting fewer pages always produces a proportionally smaller file.
Running several downloads concurrently
Create one ClientSession and reuse it for multiple URLs. A bounded semaphore prevents an untrusted batch from opening unlimited simultaneous transfers:
semaphore = asyncio.Semaphore(5)
async def bounded_download(session, url, destination):
async with semaphore:
await download_pdf(session, url, destination)
Give each output a unique, trusted path. If URLs or filenames come from users, validate schemes, destinations, and allowed hosts to reduce server-side request forgery and path-traversal risk. The appropriate policy depends on your application.
Free tools Windows power users keep installed
One-click scans. No signup required.
Verify the result
After writing, reopen the output and check its page count. This catches mistakes in selection and confirms that the writer produced a readable document:
result = PdfReader("selected-pages.pdf")
assert len(result.pages) == len(selected_pages)
For high-value workflows, also compare the intended order and inspect representative pages visually. A syntactically readable PDF can still contain a page with missing external content or an unexpected crop box.
Or skip the browser setup
If your wider workflow starts with capturing a web page rather than downloading an existing PDF, ScreenshotNeo provides a single HTTP request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', body);
See the ScreenshotNeo documentation for options such as PDF output, full-page capture, custom headers, cookies, JavaScript, waiting conditions, and signed webhooks. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Can aiohttp select PDF pages by itself?
No. It transfers HTTP responses; use a PDF library such as pypdf for page selection and writing.
Best Value
Should I keep the downloaded source file?
Keep it when you need auditability, retries, or diagnosis. Otherwise, a temporary file can be removed after the output has been verified.
Does selecting pages preserve links and annotations?
Preservation depends on the source PDF and the objects referenced by each page. Verify important interactive content in the generated file.
Frequently Asked Questions
Can aiohttp select PDF pages by itself?
No. It transfers HTTP responses; use a PDF library such as pypdf for page selection and writing.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchShould I keep the downloaded source file?
Keep it when you need auditability, retries, or diagnosis. Otherwise, a temporary file can be removed after the output has been verified.
Does selecting pages preserve links and annotations?
Preservation depends on the source PDF and the objects referenced by each page. Verify important interactive content in the generated file.
The Bottom Line
Download with aiohttp, stream large responses to disk, convert human page numbers to zero-based indexes, validate them, and let pypdf write the selected pages.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

