To download PDFs linked from a website, use GNU Wget’s recursive mode with a PDF suffix filter. For example, wget --recursive --level=inf --no-parent --accept=pdf https://example.com/documents/ follows links it can discover from that starting point and saves matching URLs. It cannot guarantee every PDF on a site: documents not linked along the crawl path, files with no .pdf suffix, and content behind inaccessible or interactive pages may be missed.
Download linked PDFs with GNU Wget
Wget is suitable when the site exposes document links in ordinary pages that it can reach from a starting URL. Its recursive retrieval follows links found in HTML and CSS, including href, src, and CSS url() references; it traverses breadth-first and lets you set a depth limit. The crawl is therefore a traversal of a discoverable link graph, not a complete inventory of a domain. See the GNU Wget manual’s recursive download documentation.
Basic command
Install GNU Wget for your operating system if it is not already available, then open a terminal and substitute the page or directory that links to the PDFs:
wget --recursive --level=inf --no-parent --accept=pdf https://example.com/documents/
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
--recursiveenables link-following retrieval.--level=infremoves the finite recursion-depth limit. It does not make pages undiscoverable to Wget appear.--no-parentprevents ascent to parent directories in the URL hierarchy. It is a useful boundary when the starting URL is in a directory, but it is not a universal domain-wide scope control.--accept=pdfaccepts URLs whose names match thepdfsuffix filter. Wget’s accept and reject rules match names, suffixes, and patterns; they do not establish that the downloaded response content is actually a PDF. See the Wget manual’s file-type filtering documentation.
Run the command from the directory where you want the downloaded site tree created. Wget generally recreates paths under a directory named for the host. Review that output before moving or renaming files so that similarly named documents from different paths do not overwrite one another.
Limit the crawl depth
An unlimited depth can follow a large link graph. If you want only pages close to your starting point, replace inf with a number, such as --level=2. This caps the number of link levels Wget follows; PDFs linked beyond that limit may not be found. Choose depth based on how the site organizes its document links rather than assuming a shallow crawl is complete.
Exclude unwanted paths or file patterns
Wget can reject suffixes or patterns as well as accept them. For example, to accept PDF paths while rejecting a known section, use --reject-regex only when you have a pattern that matches the unwanted URLs. Pattern rules need to fit the actual URL form; inspect the site structure and Wget output rather than relying on an imagined directory layout. The file-type filtering documentation describes accept/reject behavior.
If a site serves a PDF at a URL without a .pdf ending, a suffix filter can omit it. Conversely, a URL ending in .pdf can return an error page or other content. Verify important downloads by opening them or checking their response/content independently.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsRank #2
What “all PDFs” can and cannot mean
A recursive download can retrieve only resources its crawl reaches and accepts. Starting from a single page does not expose a site’s private file store or automatically enumerate every URL. A PDF may be missed if it is not linked from a reachable page, if its URL is generated only by client-side behavior Wget does not execute, if the link requires a session, or if the URL does not match the suffix filter. These are limits of link discovery and request access, not proof that the file does not exist.
- Start-point coverage: begin at the page, index, or section where documents are linked. Multiple unrelated sections may require separate runs.
- URL coverage: links can lead off-site or into large parts of a domain. Set a sensible boundary and review the URLs being fetched.
- File identification: suffix filtering is convenient for conventional
.pdflinks, but it is not content validation. - Access: a crawler cannot download material that requires credentials or access rights it does not have. Do not attempt to bypass protections.
Respect crawl policy and access rights
GNU Wget documents that recursive retrieval respects the Robot Exclusion Standard; see its robots.txt documentation. Google explains that robots.txt gives crawler instructions and helps manage crawling, but is not a security mechanism: a disallowed URL may still be indexed if linked elsewhere, and crawl blocking can affect documents such as PDFs. See Google Search Central’s robots.txt guide.
For a third-party site, keep downloads within the scope the site permits and avoid treating a publicly reachable URL or a robots.txt rule as permission to access protected material. Site owners who need confidentiality should use actual access controls, not robots.txt alone.
Use Scrapy when you need a controlled crawler
For a developer workflow with custom URL discovery, file naming, or storage, Scrapy’s Files Pipeline can download file URLs supplied in an item, write them to configured storage, and return success or failure details. It is a crawler framework rather than a one-click bulk downloader. Consult the official Scrapy media pipeline documentation for setup and supported settings.
Recommended Free Tools
Rank #3
Workflow at a glance
- Build a spider that visits the permitted pages and collects the PDF URLs you want to download.
- Put those URLs into the item field configured for Scrapy’s Files Pipeline.
- Set
FILES_STOREto the destination storage location and enable the pipeline in the project settings. - Optionally customize
file_pathto control names and folder structure. - Inspect each pipeline result for success or failure and handle failed URLs deliberately.
Scrapy does not remove the need to solve discovery: the spider still has to find the relevant links and obey the site’s crawl policy. Its advantage is control over collection and storage when Wget’s recursive traversal and suffix filter are not enough.
Choose the method that fits the job
| Need | Better fit | Trade-off |
|---|---|---|
| Download conventional PDF links reachable from a starting page or section | GNU Wget recursive mode with an accept suffix | Easy to run, but limited to links it discovers and URLs matching the filter. |
| Control URL discovery, file naming, storage, and per-file outcomes | Scrapy Files Pipeline | Requires building and configuring a crawler. |
| Obtain every document, including unlinked or protected files | No generic crawler can guarantee this | Ask the site owner for an official archive or authorized export. |
Troubleshooting missed or failed PDFs
The command finishes but finds no PDFs
Check that the start URL actually links to documents, and that the links are reachable in pages Wget can retrieve. Confirm that the URLs end in .pdf if you used --accept=pdf. Try the relevant section’s index page as the starting point; avoid assuming that a homepage links to every document area.
Some documents are missing
Look for links beyond the chosen --level, URL paths outside the starting directory, and PDF links without a conventional suffix. If the site needs a login, a browser-only interaction, or a script-generated link, Wget’s recursive link following may not find or access it. For a permitted workflow requiring custom extraction, use a crawler such as Scrapy to collect URLs explicitly.
The crawl downloads too much
Reduce recursion depth, choose a narrower starting URL, and add a suitable reject pattern for irrelevant paths. Check the site’s crawler policy before broadening scope. A recursive crawl without a carefully chosen boundary can traverse many pages that are unrelated to the documents you need.
A downloaded “PDF” will not open
The URL suffix filter is based on the URL name, not a guarantee about response contents. The server may have returned an error page, login page, or another resource under a PDF-looking URL. Check the response and access requirements, then verify the file itself rather than trusting its filename.
Wget does not follow a link seen in a browser
The page may depend on client-side JavaScript, a user session, or an interaction to expose the URL. Wget’s documented recursive behavior follows links it finds in retrieved HTML and CSS; it is not a general browser automation engine. For authorized access, use the site’s export/download feature or a crawler adapted to its documented structure.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your next step is to capture a webpage as an image or PDF rather than collect PDF files linked across a site, ScreenshotNeo offers a website screenshot API. A single request can return a screenshot or PDF; cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed, and an MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. This is a page-capture service, not a bulk PDF-link crawler.
cURL example, with the ScreenshotNeo API documentation:
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Best Value
Learn more at ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.
Frequently Asked Questions
Can Wget download PDFs from an entire website?
Only PDFs discoverable from the starting URL within the crawl’s scope and accepted URL patterns; it cannot guarantee a complete site inventory.
Does robots.txt make a PDF private?
No. Google describes robots.txt as crawler guidance, not access control. Use authentication or other real access controls for confidential files.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

