Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If a PDF downloaded with Python is only about 1 KB, the file size alone does not reveal the cause. The server may have returned an error or access page instead of a PDF, or the transfer may have stopped early. Check the HTTP response and the bytes you saved before changing libraries or adding request headers.

Check what the server actually returned

A filename ending in .pdf does not prove that the response contains a PDF. Start by checking the response status, final URL, headers, and the beginning of the saved content. An HTTP response object or a file on disk is not, by itself, proof that the request succeeded; Requests recommends checking status_code or calling raise_for_status() (Requests Quickstart).

  • Status: A failed status can identify an HTTP error. Call raise_for_status() before writing the body so an unsuccessful response raises an exception.
  • Final URL: Check response.url to see where the request ended after any redirects.
  • Headers: Inspect Content-Type to see the server’s claimed content type and Content-Length, if present, for a declared size. These headers are clues, not proof that the content is a complete, valid PDF.
  • Leading bytes: Read a short sample from the saved file. If it contains readable text or markup, the response may be an error, login, or access page rather than the document you wanted.

The reported size of about 1 KB cannot distinguish those cases from an incomplete transfer or another server response. The URL, response, and file contents are needed to identify the actual cause.

Download with Requests using a binary file and streamed chunks

Requests documents iter_content() for writing streamed downloads to a file opened in binary mode. This example also prints useful response details and checks the HTTP status before saving the body:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

url = "https://example.com/file.pdf"

with requests.get(url, stream=True, timeout=30) as response:
    response.raise_for_status()
    print("Final URL:", response.url)
    print("Status:", response.status_code)
    print("Content-Type:", response.headers.get("Content-Type"))
    print("Content-Length:", response.headers.get("Content-Length"))

    with open("download.pdf", "wb") as output:
        for chunk in response.iter_content(chunk_size=64 * 1024):
            if chunk:
                output.write(chunk)

Replace the example URL with the document URL. The wb mode writes bytes without text encoding or newline conversion. iter_content() yields chunks and handles gzip and deflate transfer encodings; Requests distinguishes it from Response.raw, which exposes the raw stream (Requests Quickstart).

With stream=True, the response connection remains open until the body is consumed or the response is closed. The with block closes it even if the body is not fully read. Consume the body as in the example, or close the response when stopping early (Requests Advanced Usage).

Check whether the download is incomplete

If the server provides a usable Content-Length, compare its declared byte count with the number of bytes written. A mismatch is evidence of a short read, but a matching count does not prove that the body is a valid PDF: the server could have sent a complete error page, or its length header could be wrong. If there is no Content-Length, this simple size comparison is unavailable.

A historical Requests issue opened on June 27, 2019, described a response of 2,583 bytes against a declared Content-Length of 66,892,906. The issue illustrates a length mismatch and is not evidence that it explains this particular 1 KB file or a guarantee about current Requests behavior (Requests issue #5124).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the next step from the response

What you find What to investigate
An error status, or readable error text in the file Read the status and response body to identify the server’s error. Do not treat the saved response as the PDF.
A login or access page Check whether the URL requires legitimate authentication, a session, or permission to access the document. Adding a browser User-Agent is not a reliable general fix.
A different final URL Check whether the redirect led to the intended document or to an access, login, or other endpoint.
A declared length larger than the bytes written Investigate an interrupted transfer or a server-side length mismatch; retry only after checking that the URL and response are appropriate.
Bytes and headers that appear consistent, but the file still will not open Check the leading bytes and the response’s claimed type. Size and headers alone cannot establish that the content is a valid PDF.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When to use Python’s urlretrieve()

The standard-library helper urllib.request.urlretrieve() copies a URL resource to a local file. Python 3.13.15 documents that it raises ContentTooShortError when it detects fewer downloaded bytes than the amount reported by Content-Length. If the header is absent, it cannot check the downloaded size and simply returns the file. That check can catch some short reads, but it does not establish that the file is a valid PDF or that the server’s declared length is correct (Python 3.13.15 urllib.request documentation).

Option Useful when Documented limitation
Requests with iter_content() You need streamed chunks and access to response details such as status, final URL, and headers. You must check the response and ensure a streamed body is consumed or the response is closed.
urllib.request.urlretrieve() You want a direct standard-library URL-to-file helper. It detects a short read only against a supplied Content-Length; it does not validate PDF content.

Switching libraries alone does not identify why a particular URL produced a small file. First establish what the server returned and whether the saved bytes match the expected response.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.