JavaScript can help you download a website, but it does not make a complete offline mirror automatic. If the pages and links are already in HTML, a recursive downloader such as GNU Wget is usually the more direct way to retrieve pages and assets and rewrite links for local browsing. If a site creates important content or links only after JavaScript runs, use browser automation such as Playwright to render pages and discover those links. A parser such as Cheerio reads markup; it does not run the page’s JavaScript.
The example below uses Playwright to crawl a deliberately bounded set of rendered pages and save their rendered HTML. It is a starting point for a JavaScript-heavy site, not a guarantee of a self-contained or complete copy: remote images, stylesheets, scripts, login state, and other dependencies may still need separate handling.
Choose the kind of offline copy you need
“Download an entire website” can mean anything from saving a few pages for reference to creating a locally navigable mirror with images, stylesheets, scripts, and documents. Decide what counts as in scope before crawling. Set the allowed host or hosts, paths to include or exclude, a finite page limit, and which assets matter. A website can also depend on APIs, authentication, or behavior that cannot be reproduced just by saving its pages.
| Situation | Good starting point | Important limitation |
|---|---|---|
| Links and page content are present in ordinary HTML, XHTML, or CSS | Recursive retrieval with Wget | It can follow links and convert downloaded links for local browsing, but you still need to bound and verify the result. |
| Important content or links appear only after client-side JavaScript runs | Playwright or another browser automation tool | A browser can render the page, but Playwright’s download-event API is for page-triggered file downloads, not a turnkey full-site mirroring feature. |
| You already have HTML and want to inspect or transform it | Cheerio | Cheerio parses markup; it does not execute scripts, render CSS, or load external resources. |
GNU Wget documents recursive retrieval, reconstruction of the remote directory structure, and conversion of links for local browsing. That makes it a sensible first choice when a conventional mirror is the goal and the required links are discoverable in markup. A JavaScript crawler is useful when rendering is necessary to reveal content or navigation, but it does not remove the need to collect dependent files and test the local result.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
- Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Prepare a safe, bounded crawl
- Confirm permission and intended use. Check the site’s terms and any applicable permission requirements, especially before redistributing a copy. The rules for a particular site cannot be inferred from a generic crawler guide.
- Inspect robots.txt and follow applicable crawler directions. Wget says it respects the Robot Exclusion Standard. Google’s guidance explains that robots.txt is primarily for managing crawler access and traffic, not for protecting private files or keeping a page out of search. It is crawler guidance, not a security boundary or permission grant.
- Limit the crawl. Keep it to an allowed host and relevant paths. Search pages, calendars, filters, and query parameters can generate very large numbers of distinct URLs. Exclude or normalize them where appropriate, and use a maximum page count.
- Choose based on where content is generated. If the links and content are already in markup, a recursive downloader may suffice. If they appear only after JavaScript runs, render pages in a browser and collect links from the rendered page.
- Verify locally. Open representative saved pages and check navigation, images, styles, scripts, and documents. Treat the output as a bounded offline copy, not a perfect clone.
Use Playwright to crawl rendered pages with JavaScript
This Node.js example opens each page in a real browser, waits for DOM parsing, collects same-host links from the rendered DOM, and saves each page’s rendered HTML in a local folder. It stops at a page cap, avoids revisiting URLs, and skips links with query strings or fragments to reduce duplicate and unbounded crawling. Change the starting URL and scope checks for your site before running it.
Install the prerequisites
Install a current Node.js release, create a project directory, and install Playwright:
npm init -ynpm install playwrightnpx playwright install chromium
Playwright’s browser installation may download browser binaries. If your environment uses a proxy or restricts downloads, configure that environment according to your organization’s network policy.
Rank #2
- Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Save this crawler as crawl.mjs
import { chromium } from 'playwright';
import fs from 'node:fs/promises';
import path from 'node:path';
const startUrl = new URL('https://example.com/');
const allowedHost = startUrl.host;
const maxPages = 100;
const outputDir = path.resolve('site-copy');
function outputPath(url) {
let pathname = decodeURIComponent(url.pathname);
if (pathname.endsWith('/')) pathname += 'index.html';
else if (!path.extname(pathname)) pathname += '/index.html';
return path.join(outputDir, url.host, pathname.replace(/^/+/, ''));
}
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();
const queue = [startUrl.href];
const visited = new Set();
try {
while (queue.length && visited.size < maxPages) {
const current = new URL(queue.shift());
current.hash = '';
if (current.host !== allowedHost || current.search) continue;
const key = current.href;
if (visited.has(key)) continue;
visited.add(key);
const page = await context.newPage();
try {
const response = await page.goto(key, {
waitUntil: 'domcontentloaded',
timeout: 30000
});
if (!response || !response.ok()) {
console.warn(`Skipping ${key}: HTTP ${response?.status() ?? 'no response'}`);
continue;
}
const html = await page.content();
const destination = outputPath(current);
await fs.mkdir(path.dirname(destination), { recursive: true });
await fs.writeFile(destination, html, 'utf8');
console.log(`Saved ${key} -> ${destination}`);
const links = await page.locator('a[href]').evaluateAll(anchors =>
anchors.map(anchor => anchor.href)
);
for (const href of links) {
try {
const next = new URL(href);
next.hash = '';
if (next.host === allowedHost && !next.search && !visited.has(next.href)) {
queue.push(next.href);
}
} catch {
// Ignore links that are not valid absolute URLs.
}
}
} catch (error) {
console.warn(`Failed ${key}: ${error.message}`);
} finally {
await page.close();
}
}
console.log(`Finished: visited ${visited.size} page(s).`);
} finally {
await context.close();
await browser.close();
}
The example saves each page’s browser-rendered DOM. It does not download and rewrite every image, stylesheet, font, script, or document reference. Many saved pages will continue to request those resources from the original site, so this output alone may not work offline. Nor does the example handle login, infinite scrolling, click-to-reveal interfaces, or URL patterns that require query parameters. Add only the behavior and scope you need, and retain the page cap.
Run it and inspect the result
- Replace
https://example.com/with the authorized starting URL. - Adjust
allowedHost,maxPages, and the URL filters to match the intended scope. For subdomains, define an explicit allowlist rather than accepting every host. - Run
node crawl.mjs. The saved HTML files appear undersite-copy. - Open several files locally. Check that the expected text is present and test whether navigation and assets work without relying on the original site.
When Wget is a better fit
If the site’s links are discoverable in ordinary markup and the goal is a conventional offline mirror, use Wget’s documented recursive retrieval and link-conversion capabilities rather than writing a JavaScript crawler solely to fetch HTML. Its manual describes options for recursive retrieval, directory structure, and converting links for local viewing. Review those options for your installed Wget version and set explicit scope and depth limits; do not assume one command will correctly capture every application.
Wget can follow HTML, XHTML, and CSS links, but it cannot make server-dependent behavior or JavaScript-created content magically available. A useful division of labor is to use browser automation when rendering reveals URLs that static retrieval would miss, then use a recursive downloader where appropriate for linked resources. Validate the resulting files rather than treating a successful command exit as proof of completeness.
Rank #3
- High capacity in a small enclosure – The small, lightweight design offers up to 6TB* capacity, making WD Elements portable hard drives the ideal companion for consumers on the go.
- Plug-and-play expandability
- Vast capacities up to 6TB[1] to store your photos, videos, music, important documents and more
- SuperSpeed USB 3.2 Gen 1 (5Gbps)
Why pages or assets may be missing
- The page is rendered after navigation. A parser such as Cheerio does not execute JavaScript, and capturing too early in a browser can miss content added later. Wait for a specific selector or a suitable page condition where needed; avoid assuming one fixed delay works on every site.
- Links require interaction. Menus, “load more” buttons, and scrolling may reveal links only after user actions. The sample follows links already present in the rendered DOM; it does not click controls or scroll through an infinite list.
- URLs were excluded by the scope filter. The example skips all query-string URLs. That prevents some duplicate or unbounded paths, but it will also omit legitimate pages whose content depends on query parameters. Add a narrow, deliberate rule if those pages are required.
- The saved HTML depends on remote assets.
page.content()captures markup, not the referenced files. A saved file may look unstyled or have missing images when offline. Downloading assets and rewriting references is a separate part of building a self-contained mirror. - A request failed or returned an error. The sample logs non-success HTTP responses and navigation errors, then continues. Review those messages and determine whether the cause is a transient failure, a blocked request, an unavailable page, or a site access rule; do not simply increase request volume.
- Pages need a session or special access. The sample uses a fresh browser context and no credentials. If the intended material requires authentication, use only an authorized account and handle any session information securely.
Performance, reliability, and responsible use
The script processes pages sequentially and has a configurable page cap; it does not publish a speed or completeness guarantee. Raising the cap or running many crawls in parallel increases requests and local storage, and can burden a site. Keep the crawl as small as the use case permits, respect site guidance, and stop if the site responds with restrictions or errors. Browser automation also requires browser binaries and more resources than parsing a saved HTML document.
Robots.txt should not be treated as protection for sensitive content. Google’s Search Central documentation states: “This is used mainly to avoid overloading your site with requests; it is not a mechanism for keeping a web page out of Google.” Protect private material with access controls, and do not use a crawler to bypass them.
Or skip the browser setup
If you only need a screenshot of a page rather than a locally navigable website, ScreenshotNeo is a website screenshot API and MCP server for developers. It is not a full-site downloader and does not create an offline website mirror. One API call captures a page as an image or PDF; its documented options include full-page capture and browser-related settings.
Rank #4
- Plug-and-play expandability
- SuperSpeed USB 3.2 Gen 1 (5Gbps)
For example, this cURL request captures a screenshot of the target URL. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
- Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
- An MCP server provides screenshot tools for AI agents, including Claude, Cursor, and other MCP clients.
- The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan to try up to 1,000 screenshots a month with no card.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Frequently Asked Questions
Can JavaScript download all pages from a website?
Not by itself in a dependable, universal way. A script can crawl a bounded set of pages, but the site’s links, access rules, rendering behavior, and assets determine what can be saved.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Best Value
- 【Upgraded version】 - The mirror logo strip is combined with the striped non-slip design. The rounded corners of the shell are more suitable for holding. The strips play a heat dissipation function to ensure a stable and fast transmission process.
- 【Ultra-thin and quiet】 - The motherboard adopts JMicron 578 noise-free solution, giving you a quiet working environment. Lightweight and portable size designed to fit in your pocket for easy portability.
- 【Ultra-Fast Data Transfers】 - Pairing this external hard drive with JMicron 578 solution USB 3.0 and USB 2.0 interfaces enables blazing-fast data transfer. It boasts theoretical read speeds of up to 125MB/s and write speeds of up to 103MB/s.
- 【Plug and Play】 - With no software to install, just plug it in and the drive is ready to use.The hard disk chip is wrapped with an aluminum anti-interference layer to increase heat dissipation and protect data.
- 【What You Get】 - 1 x Portable Hard Drive, 1 x USB 3.0 Cable, 1 x User Manual, Gift-type shell packaging ,Three-year manufacturer's warranty and free technical support services.
Does the Playwright example create a fully offline website?
No. It saves rendered HTML pages, but does not fetch and rewrite every linked resource. Test the local files and add asset handling if a self-contained copy is required.
Can robots.txt authorize me to copy a site?
No. It provides crawler guidance, not permission to republish content or a security boundary.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

