Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallFor a small or one-off PDF download, Python’s built-in urllib.request.urlopen() can fetch the response and save its bytes to a file. For large files, use Requests with stream=True and write chunks as they arrive. In both cases, set a timeout, handle HTTP errors, and remember that a URL ending in .pdf does not prove the response is a PDF.
Download a PDF with Python’s standard library
This short-file example needs no third-party package. It reads the response as bytes and writes them in binary mode (wb), which is appropriate for PDF data. The timeout is an example value; choose one that suits the server and your application.
from pathlib import Path
from urllib.request import urlopen
url = "https://example.com/document.pdf"
out = Path("document.pdf")
a— = 1
with urlopen(url, timeout=30) as response:
out.write_bytes(response.read())
print(f"Saved {out}")
Replace the example URL with the PDF’s actual URL and choose the output path you want. The destination is relative to the program’s current working directory unless you give it an absolute path. If a file already exists at that path, this example replaces it.
urlopen() returns a context-manager response; its body is bytes, and the timeout bounds how long the connection can wait. The Python 3.13 documentation also describes redirects and authentication support in urllib.request. For a large response, avoid response.read() because it holds the entire body in memory. The Python documentation recommends Requests as a higher-level HTTP interface, but the standard library is enough for a straightforward download.
Recommended Free Tools
#1 Best Overall
See the Python 3.13 urllib.request documentation for the module’s URL-opening functions and exceptions.
Stream a large PDF with Requests
Requests downloads response content immediately by default. For a larger file, pass stream=True, check that the HTTP response succeeded, and write non-empty chunks to disk. This keeps the whole response from being loaded into memory at once.
from pathlib import Path
import requests
url = "https://example.com/document.pdf"
out = Path("document.pdf")
with requests.get(url, stream=True, timeout=(5, 60)) as response:
response.raise_for_status()
with out.open("wb") as file:
for chunk in response.iter_content(chunk_size=1024 * 64):
if chunk:
file.write(chunk)
print(f"Saved {out}")
The tuple sets example connect and read timeouts, and the chunk size is an example—not a universal performance setting. Adjust them for your network and expected file sizes. The with blocks close both the response and the file if an exception interrupts the download.
Rank #2
Install Requests if it is not already available in your Python environment:
python -m pip install requests
Requests documents raise_for_status() for unsuccessful HTTP responses and recommends iter_content() for streamed file saving. Use iter_content() for ordinary downloads rather than reading the response’s raw stream directly. See the Requests Quickstart and Requests Advanced Usage.
Choose between urllib and Requests
| Need | Use | Why |
|---|---|---|
| No extra dependency | urllib.request |
It is part of Python and can open URLs and return response bytes. |
| Simple status handling | Requests | raise_for_status() makes it straightforward to reject unsuccessful HTTP responses. |
| Large file written incrementally | Requests with stream=True and iter_content() |
Chunks can be written as they arrive instead of retaining the complete body in memory. |
| One small, uncomplicated download | Either | Use urlopen() for a built-in solution, or Requests if your project already uses it. |
The Python standard-library interface also raises URLError for URL-related failures; an HTTP failure can be an HTTPError, which is a subclass of URLError. With Requests, call raise_for_status() or inspect status_code before treating the response as a successful download.
Handle redirects, errors, and unexpected content
A URL can redirect to another address or return a login page, an access-denied page, or an HTML error body rather than a PDF. A .pdf suffix is only a naming clue: servers can return different content at that address. Check the HTTP result, then validate the saved data if your application depends on receiving a PDF.
Check the file signature when a PDF is required
One practical sanity check is to inspect the beginning of the downloaded file for the PDF header marker %PDF-. For example, after writing a file:
with out.open("rb") as file:
header = file.read(5)
if header != b"%PDF-":
raise ValueError("The downloaded response does not start with a PDF header")
This catches many accidental HTML or text responses, but it is not a complete PDF validator. If validity is important, use a PDF-aware parser or validator in your application as a separate step. The HTTP client documentation explains responses and status handling; it does not prescribe a PDF-signature validation method.
Respect authentication and access controls
Some documents are available only after authentication or with required cookies or headers. Use credentials and access methods you are authorized to use; do not treat a download script as a way around a site’s access controls. Both the standard library and Requests provide ways to send request headers or authentication information, but the exact requirements depend on the service.
Make the download safer for applications
The basic examples write directly to the destination. If interruption must not leave a partial file under the final filename, write to a temporary path and rename it only after the response has been fully read and any checks have passed. Also decide deliberately whether an existing destination should be replaced, rejected, or versioned; there is no single overwrite policy that fits every application.
For repeated downloads, useful safeguards include:
- Choose a deliberate output directory and ensure the process has permission to write there.
- Use a timeout suited to the server and expected transfer duration; a timeout that is too short can interrupt a slow but valid download, while no timeout can leave a process waiting indefinitely.
- For streamed Requests responses, consume the body or close the response. The context manager shown above handles cleanup if the body is only partly read, freeing the connection for reuse.
- Log the source URL, destination, and exception details without logging secrets embedded in URLs or credentials.
- Do not trust the filename extension or a successful HTTP status alone when downstream steps require a valid PDF.
Troubleshooting common download failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Requests raises an HTTP error | The server returned an unsuccessful status, such as an access or not-found response. | Check the URL and permissions. Inspect the response status and, if appropriate, the response headers or body to understand what the server returned. |
| The saved file opens as HTML or is not a PDF | The URL redirected or served a login, access-denied, or error page. | Check HTTP success and inspect the downloaded content. Supply authorized authentication details if the document requires them. |
| The request times out | The server or network did not respond within the configured timeout. | Check connectivity and the URL, then choose timeout values appropriate for the server and file size. A longer timeout cannot fix a permanently unavailable endpoint. |
| A large download uses too much memory | The response was read all at once, for example with response.read() or Requests’ default non-streamed behavior. |
Use a streamed approach and write chunks to disk as they arrive. |
| A connection seems unavailable after an interrupted streamed request | The response body was not consumed and the response was not closed. | Use a response context manager, as in the Requests example, so a partial read still closes the response. |
| Python cannot write to the destination | The destination directory may not exist or the process may lack write permission. | Use an existing writable directory or create the intended directory before opening the output file. |
Or skip the browser setup
If what you need is a PDF rendition of a live webpage rather than an existing PDF file, ScreenshotNeo is a website screenshot API with PDF capture. It is not a replacement for downloading a PDF already hosted at a URL. The call below is the supplied one-call screenshot example; it saves a WebP image, not a PDF. See the ScreenshotNeo documentation for its PDF options and other parameters.
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
- Before capture, it accepts the cookie or consent banner as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off.
- Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing status in headers.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents, including Claude, Cursor, and other MCP clients. - The Free plan includes 1,000 screenshots per month with no card required; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
Sources
Frequently Asked Questions
Can I use Python’s urlretrieve() for this?
Yes. Python documents urlretrieve() as a way to copy a URL resource to a local file, but it is in the legacy interface section. For a new script, urlopen() makes it easier to show a timeout and explicit response handling.
Does the PDF have to be publicly accessible?
No, but a private document requires the authentication or access method its host authorizes. A script cannot download content the server does not permit the request to access.
Will this work for a URL that does not end in .pdf?
It can, if that URL actually returns PDF bytes. The filename suffix is not what determines the response format.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

