Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Cheerio parses HTML that your Node.js program already has; it does not open a browser or run page JavaScript. The practical workflow is: install Cheerio, fetch or otherwise acquire markup, call cheerio.load() (or the loader that matches your input), select elements with CSS selectors, extract text and attributes, then save structured records. If a page builds its content in the browser, render it with a browser-capable step first and give the resulting HTML to Cheerio.
What Cheerio does—and where it stops
Cheerio is a fast parser and DOM-like manipulation API for HTML and XML. Its jQuery-style API lets you query, traverse, read, change and serialize markup. It is not a web browser: it does not visually render pages, load external resources, or execute JavaScript. A server-rendered article is therefore straightforward; a product list inserted by client-side JavaScript will be absent from the HTML response that Cheerio receives.
Keep acquisition and parsing as separate responsibilities. Your HTTP client (such as Node’s built-in fetch) handles status codes, headers, cookies, redirects, retries and rate limits. Cheerio turns the resulting bytes or string into a queryable document.
Install Cheerio and import it
The current official introduction says Cheerio runs on Node.js 22.19 or later. Confirm the runtime requirement for the version you pin in production; the npm registry currently lists Cheerio 1.2.0 and MIT licensing, values that can change.
#1 Best Overall
npm install cheerio
Use ESM in a project whose package.json contains "type": "module":
import * as cheerio from 'cheerio';
For CommonJS:
const cheerio = require('cheerio');
Pin the dependency (for example, with your lockfile), verify that your deployment uses the supported Node version, and rerun installation in that same environment. An older release note mentioned Node.js 18.17 or newer, but do not assume that historical minimum applies to the current release.
Minimal static-page scraper
This complete ESM example fetches a page, checks the response, parses it, and extracts a heading and links:
import * as cheerio from 'cheerio';
const response = await fetch('https://example.com', {
headers: { 'user-agent': 'my-research-bot/1.0' }
});
if (!response.ok) {
throw new Error(`HTTP ${response.status} ${response.statusText}`);
}
const html = await response.text();
const $ = cheerio.load(html);
const title = $('h1').first().text().trim();
const links = $('a[href]').map((_, el) => ({
text: $(el).text().trim(),
href: $(el).attr('href')
})).get();
if (!title && links.length === 0) {
throw new Error('Expected selectors did not match; the page shape may have changed.');
}
console.log({ title, links });
fetch obtains the response; cheerio.load parses the returned markup. Always validate status and important selections so a redesign does not silently create empty data.
Choose the loader that matches your input
| Input | Method | When to use it |
|---|---|---|
| Already-decoded HTML string | load(markup) |
Normal text responses or locally stored markup. |
| Raw bytes | loadBuffer(buffer) |
Encoding is uncertain; Cheerio can perform byte-oriented encoding detection. |
| Text stream | stringStream() |
The stream has already been decoded to text. |
| Byte stream | decodeStream() |
Keep decoding and parsing in the stream pipeline. |
| URL | fromURL(url) |
Let Cheerio perform the fetch when you do not need custom HTTP policy. |
Only load is included in the browser build. Explicit fetching is usually easier to operate because your code can set headers, enforce timeouts, inspect status, retry selected failures and obey a site’s rate limits.
Select, traverse and extract safely
Core CSS selectors
Cheerio’s selector engine supports tag, class, ID, attribute, universal and documented pseudo-class selectors. Prefer stable semantic attributes over brittle chains of positional selectors.
Rank #2
const $ = cheerio.load(html);
const heading = $('h1').first().text().trim();
const firstCard = $('.card').first();
const cardTitle = firstCard.find('.title').text().trim();
const href = firstCard.find('a').attr('href');
const prices = $('[data-price]').map((_, el) => $(el).attr('data-price')).get();
.text() combines descendant text; .attr('href') reads an attribute and returns undefined when it is missing. Use .first() when you intentionally need one match, and check .length before treating a required element as present.
Extract repeatable records with extract
For lists of cards, products, articles or links, define the output shape once:
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsconst records = $.extract({
articles: [{
selector: 'article',
value: {
title: 'h2',
summary: '.summary',
url: { selector: 'a', value: 'href' }
}
}]
});
console.log(records.articles);
A selector string returns the first matching text value. Object descriptors can read attributes or properties such as outerHTML, innerHTML, tagName and innerText. This keeps a scraper’s schema explicit and makes it easier to test when a site changes.
Normalize URLs and whitespace
const absoluteUrl = (href, pageUrl) => {
try { return new URL(href, pageUrl).href; }
catch { return null; }
};
const rows = $('article').map((_, el) => {
const link = $(el).find('a[href]').attr('href');
return {
title: $(el).find('h2').text().replace(/\s+/g, ' ').trim(),
url: link ? absoluteUrl(link, 'https://example.com/news') : null
};
}).get();
Resolve relative links against the page URL, preserve missing values as null where your schema permits them, and log malformed URLs instead of dropping records invisibly.
Parse fragments and serialize markup
Document mode may add html, head and body around a fragment. Pass false as the third argument when you need fragment parsing:
const $ = cheerio.load('<li>One</li>', null, false);
const fragment = $.html();
console.log(fragment);
Use $.html() to serialize the document or a selected node when you need cleaned or transformed markup. Remember that serialization reflects parser behavior; it is not a screenshot or a browser’s painted output.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
Scraping JavaScript-rendered pages
Cheerio alone cannot scrape content that exists only after client-side JavaScript runs. First acquire rendered HTML with browser automation or another DOM-emulation layer, wait for the application state you need, then pass the resulting HTML to Cheerio for fast extraction. If the data is available from a documented JSON endpoint, calling that endpoint directly can be simpler, subject to its terms and authentication requirements.
When a browser step is justified
- The initial response contains an empty application shell.
- Rows appear only after scrolling, clicking, or waiting for network activity.
- Cookies, consent choices or a logged-in session change the DOM.
- You must capture the final rendered state rather than source markup.
Keep the browser phase narrow: navigate, perform required actions, wait for a specific selector, obtain page.content(), close the browser, and let Cheerio handle repeatable parsing. This reduces resource use and keeps selectors and record mapping in one place.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. A single request can capture a rendered page as PNG, JPEG, WebP or PDF; it accepts the cookie or consent banner before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status.
For a quick rendered artifact, use the documented API pattern:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for all options. It also offers custom CSS and JavaScript, selector waits, network-idle waits, device presets, full-page capture with lazy images loaded, cookies, headers, geolocation, dark mode, PDFs, signed links, asynchronous webhooks, bulk capture and a usage API. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.
Parser configuration: parse5 or htmlparser2
Cheerio uses parse5 by default, providing standards-oriented parsing and browser-like error correction. htmlparser2 is an option when more forgiving parsing or lower memory use suits a particular input, including XML-like markup. Its error correction can differ from browser standards, so test the output against representative malformed documents before switching.
Rank #4
Treat parser choice as an explicit trade-off: parse5 for standards fidelity; htmlparser2 when input characteristics or memory pressure justify it. Keep a fixture suite so a parser change cannot silently alter fields.
Recommended Free Tools
Production reliability and performance
- Bound the network: set request timeouts, cap response sizes where practical, retry transient failures with backoff, and honor robots rules, terms and rate limits.
- Validate contracts: assert required selectors, record counts and types; emit the URL and parser error when validation fails.
- Control memory: parse only what you need, release large documents after extraction, and consider byte or stream loaders for large responses.
- Cache deliberately: avoid refetching unchanged pages, but attach a freshness policy to stored records.
- Make jobs resumable: persist the last successful URL or cursor and record HTTP status, retry count and extraction version.
- Protect data: treat fetched HTML and extracted text as untrusted input; sanitize before inserting into an HTML page and keep credentials out of logs.
Cheerio’s markup-only model is generally much lighter than maintaining a browser, but it cannot replace a browser when JavaScript execution or visual state is the requirement.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common failures
“Cannot find module” or import errors
Check that npm install cheerio ran in the application directory, that your module type matches the import syntax, and that the deployment Node version satisfies the Cheerio version you installed.
Selectors return an empty string
Log the first part of the fetched HTML and inspect the response status and content type. You may have received a redirect, an access page, an empty app shell, or a redesigned DOM. Add a selector assertion instead of accepting empty records.
The page works in a browser but data is missing
That is the browser boundary: JavaScript likely creates the content after load. Use a browser-capable acquisition step, an appropriate data endpoint, or a rendered capture service, then parse the resulting HTML.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Characters are garbled
Use loadBuffer or decodeStream when you have raw bytes and encoding is uncertain. Compare the response’s declared charset with the document’s metadata.
Malformed markup produces surprising structure
Inspect the serialized output with $.html(). parse5 and htmlparser2 correct errors differently; choose the parser deliberately and add fixtures for the malformed cases you actually receive.
Scraper suddenly produces duplicate or missing rows
Check whether a selector now matches nested containers, repeated templates or pagination placeholders. Scope selectors to the item container, use .first() only when intentional, and record the source URL and extraction timestamp for diagnosis.
Choosing Cheerio versus a browser tool
| Question | Cheerio | Browser automation |
|---|---|---|
| Executes JavaScript? | No | Yes |
| Input | Strings, bytes, streams or URL | Pages plus browser context |
| Parsing | parse5 default; htmlparser2 option | Browser DOM and rendering engine |
| Resource profile | Lightweight markup parsing | Heavier browser processes and page resources |
| Best fit | Static HTML, XML and post-render extraction | Client-rendered, interactive or visual tasks |
A common architecture combines them: browser for acquisition, Cheerio for deterministic extraction and transformation.
Free tools Windows power users keep installed
One-click scans. No signup required.
Complete CommonJS example
const cheerio = require('cheerio');
async function scrape(url) {
const response = await fetch(url, { signal: AbortSignal.timeout(30000) });
if (!response.ok) throw new Error(`${url}: HTTP ${response.status}`);
const $ = cheerio.load(await response.text());
return $.extract({
title: 'h1',
links: [{ selector: 'a[href]', value: { text: 'innerText', href: 'href' } }]
});
}
scrape('https://example.com').then(console.log).catch(console.error);
Frequently Asked Questions
Does Cheerio obey robots.txt automatically?
No. Your acquisition code must implement robots, terms-of-service, authentication, rate-limit and privacy requirements appropriate to the site.
Can I use Cheerio in a browser bundle?
The documented browser build includes load, not the byte-oriented loaders. Check your bundler and target before moving a Node scraper into client-side code.
Should I scrape HTML or an API response?
Prefer a documented data endpoint when it supplies the fields you need and your use complies with its access terms; otherwise parse the appropriate HTML stage.
Is Cheerio suitable for taking screenshots?
No. Cheerio parses and serializes markup; it does not render pixels. Use a browser or a rendering API for screenshots and PDFs.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

