Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use aiohttp to download the PDF and pypdf to select and write pages. For a small file you can read the response into memory; for a large file, stream response.content to disk in chunks, then add validated, zero-based page indexes to a new PdfWriter.

What each library does

This workflow has two separate jobs:

  • aiohttp performs the asynchronous HTTP request and transfers the bytes.
  • pypdf parses the PDF, exposes its pages, and creates the selected-pages document.

aiohttp is not a PDF editor, and pypdf does not download URLs. Keeping those responsibilities separate makes failures easier to diagnose.

Install the dependencies

python -m pip install aiohttp pypdf

The examples use the current APIs represented by aiohttp 3.14.3 documentation and pypdf 6.4.2 documentation. Check the documentation for the versions pinned in your project if you are adapting the code.

Complete example: download and export selected pages

This script downloads a PDF to a temporary local path, checks the HTTP response, validates the requested indexes, and writes a second PDF.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio
from pathlib import Path
from tempfile import TemporaryDirectory

import aiohttp
from pypdf import PdfReader, PdfWriter

CHUNK_SIZE = 64 * 1024


async def download_pdf(
    session: aiohttp.ClientSession,
    url: str,
    destination: Path,
) -> None:
    async with session.get(url) as response:
        response.raise_for_status()
        with destination.open("wb") as output:
            async for chunk in response.content.iter_chunked(CHUNK_SIZE):
                output.write(chunk)


def export_pages(
    source: Path,
    destination: Path,
    page_indexes: list[int],
) -> None:
    reader = PdfReader(source)
    page_count = len(reader.pages)

    invalid = [
        index for index in page_indexes
        if index < 0 or index >= page_count
    ]
    if invalid:
        raise ValueError(
            f"Invalid page indexes {invalid}; PDF has {page_count} pages"
        )

    writer = PdfWriter()
    for index in page_indexes:
        writer.add_page(reader.pages[index])

    with destination.open("wb") as output:
        writer.write(output)


async def main() -> None:
    url = "https://example.com/document.pdf"
    selected_pages = [0, 2, 3]  # human pages 1, 3, and 4

    with TemporaryDirectory() as temporary_directory:
        source = Path(temporary_directory) / "input.pdf"
        await download_pdf(url, source)
        export_pages(source, Path("selected-pages.pdf"), selected_pages)


if __name__ == "__main__":
    asyncio.run(main())

Replace the example URL with the actual PDF URL. The resulting selected-pages.pdf contains the first, third, and fourth human-numbered pages.

Convert human page numbers to Python indexes

People normally count a PDF from page 1; Python sequences start at index 0. Convert at the boundary of your application:

  • Human page 1 becomes index 0.
  • Human pages 1, 3, and 4 become indexes [0, 2, 3].
  • Human pages 2 through 5 become indexes 1, 2, 3, 4.

For a contiguous human-inclusive range, the equivalent half-open Python slice is start - 1:end. When adding pages individually, use range(start - 1, end):

human_start, human_end = 2, 5
indexes = range(human_start - 1, human_end)
for index in indexes:
    writer.add_page(reader.pages[index])

This is Python indexing logic, not a special range syntax supplied by pypdf. Validate both endpoints before indexing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Small-download alternative: read the response into memory

For a modest PDF, an in-memory approach is shorter. aiohttp’s read(), json(), and text() convenience methods load the whole response into memory, so this version trades simplicity for memory consumption.

import asyncio
from io import BytesIO

import aiohttp
from pypdf import PdfReader, PdfWriter


async def main() -> None:
    async with aiohttp.ClientSession() as session:
        async with session.get("https://example.com/document.pdf") as response:
            response.raise_for_status()
            data = await response.read()

    reader = PdfReader(BytesIO(data))
    writer = PdfWriter()
    for index in (0, 2, 3):
        writer.add_page(reader.pages[index])

    with open("selected-pages.pdf", "wb") as output:
        writer.write(output)


asyncio.run(main())

Do not assume that streaming makes total memory usage constant: pypdf still has to parse the downloaded document, and PDF structure can require substantial memory.

Stream safely for large files

The chunked example avoids constructing one giant Python bytes object for the HTTP body. The async with blocks close the response and session, while the file context manager flushes and closes the destination.

Set practical request limits

Production code should define a timeout and, when the application accepts arbitrary URLs, enforce a maximum download size. A timeout prevents a stalled server from occupying a task indefinitely:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
timeout = aiohttp.ClientTimeout(total=90)
async with aiohttp.ClientSession(timeout=timeout) as session:
    ...

A size limit is an application policy rather than an aiohttp guarantee. Track bytes while iterating and abort when the limit is exceeded.

MAX_BYTES = 200 * 1024 * 1024
received = 0
async for chunk in response.content.iter_chunked(64 * 1024):
    received += len(chunk)
    if received > MAX_BYTES:
        raise ValueError("PDF exceeds the configured size limit")
    output.write(chunk)

Validate the response before treating it as a PDF

response.raise_for_status() converts 4xx and 5xx responses into exceptions instead of silently saving an error page. A successful status still does not prove the body is a valid PDF: servers can return HTML with status 200, redirect to a login page, or send a truncated file.

  • Inspect the final response URL when redirects matter.
  • Optionally check that the Content-Type is compatible with a PDF, while allowing servers that omit or mislabel it.
  • After download, let PdfReader parse the file and report malformed input rather than publishing the output as valid.

Handle duplicate pages, ordering, and empty selections

PdfWriter.add_page follows the order in which you call it. You can export pages in any order and can deliberately repeat an index:

for index in (4, 1, 1, 7):
    writer.add_page(reader.pages[index])

An empty selection should be rejected by your application, because writing a zero-page PDF is rarely what a user intended:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
if not page_indexes:
    raise ValueError("Select at least one page")

Common errors and fixes

401, 403, or 404 from aiohttp

The URL may require authentication, block your client, or be wrong. Confirm the URL, supply the required headers or cookies, and inspect the server’s response. Keep raise_for_status() enabled so the failure is visible.

ContentTypeError or an HTML file saved as a PDF

This often means the URL returned a login page, consent page, or error document. Follow redirects deliberately, authenticate if required, and inspect the saved bytes before passing them to pypdf.

IndexError while accessing reader.pages[index]

The requested index is negative or at least the page count. Read len(reader.pages), convert human numbers correctly, and perform explicit validation before adding pages.

pypdf cannot read the file

The input may be encrypted, malformed, truncated, or unusually complex. Preserve the original download for diagnosis, verify that the transfer completed, and handle encryption according to the document owner’s permissions. Do not promise that every PDF can be parsed without additional handling.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The process uses too much memory

Switch from await response.read() to chunked disk writing, reduce concurrency, and impose a file-size limit. Remember that parsing and writing the PDF can still consume memory.

The output is unexpectedly large

pypdf copies page objects and their referenced resources. A selected page can retain fonts, images, and other objects from the source. Avoid assuming that selecting fewer pages always produces a proportionally smaller file.

Running several downloads concurrently

Create one ClientSession and reuse it for multiple URLs. A bounded semaphore prevents an untrusted batch from opening unlimited simultaneous transfers:

semaphore = asyncio.Semaphore(5)

async def bounded_download(session, url, destination):
    async with semaphore:
        await download_pdf(session, url, destination)

Give each output a unique, trusted path. If URLs or filenames come from users, validate schemes, destinations, and allowed hosts to reduce server-side request forgery and path-traversal risk. The appropriate policy depends on your application.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Verify the result

After writing, reopen the output and check its page count. This catches mistakes in selection and confirms that the writer produced a readable document:

result = PdfReader("selected-pages.pdf")
assert len(result.pages) == len(selected_pages)

For high-value workflows, also compare the intended order and inspect representative pages visually. A syntactically readable PDF can still contain a page with missing external content or an unexpected crop box.

Or skip the browser setup

If your wider workflow starts with capturing a web page rather than downloading an existing PDF, ScreenshotNeo provides a single HTTP request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', body);

See the ScreenshotNeo documentation for options such as PDF output, full-page capture, custom headers, cookies, JavaScript, waiting conditions, and signed webhooks. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

FAQ

Can aiohttp select PDF pages by itself?

No. It transfers HTTP responses; use a PDF library such as pypdf for page selection and writing.

Should I keep the downloaded source file?

Keep it when you need auditability, retries, or diagnosis. Otherwise, a temporary file can be removed after the output has been verified.

Does selecting pages preserve links and annotations?

Preservation depends on the source PDF and the objects referenced by each page. Verify important interactive content in the generated file.

Frequently Asked Questions

Can aiohttp select PDF pages by itself?

No. It transfers HTTP responses; use a PDF library such as pypdf for page selection and writing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I keep the downloaded source file?

Keep it when you need auditability, retries, or diagnosis. Otherwise, a temporary file can be removed after the output has been verified.

Does selecting pages preserve links and annotations?

Preservation depends on the source PDF and the objects referenced by each page. Verify important interactive content in the generated file.

The Bottom Line

Download with aiohttp, stream large responses to disk, convert human page numbers to zero-based indexes, validate them, and let pypdf write the selected pages.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.