Load the HTML into Cheerio, select anchors with $('a'), and read their href attribute. Use .attr('href') for the exact string in the markup; map the selection to collect every link. If you need absolute URLs, give Cheerio a document URL and read .prop('href') instead.
import * as cheerio from 'cheerio';
const $ = cheerio.load('<a href="/docs">Docs</a><a href="https://example.com/blog">Blog</a>');
const links = $('a').map((_, el) => $(el).attr('href')).get();
console.log(links); // ['/docs', 'https://example.com/blog']
This guide shows the complete workflow: installing Cheerio, extracting one or many links, resolving relative paths, using the declarative extract API, filtering unwanted anchors, and diagnosing empty or incomplete results.
Install Cheerio and load markup
Install the package in a Node.js project:
npm install cheerio
With an ES-module project (for example, one whose package.json contains "type": "module"), import Cheerio and load a string:
import * as cheerio from 'cheerio';
const html = `
<nav>
<a href="/docs">Documentation</a>
<a href="https://example.com/blog">Blog</a>
</nav>
`;
const $ = cheerio.load(html);
In CommonJS code, use:
const cheerio = require('cheerio');
const $ = cheerio.load('<a href="/docs">Docs</a>');
Cheerio parses the markup you provide. It is not a browser and does not download a page or execute client-side JavaScript; the official introduction explains this limitation at Cheerio’s introduction.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
Get one link
Selecting anchors and calling attr('href') returns the attribute from the first matching element:
const firstHref = $('a').attr('href');
console.log(firstHref);
If the first anchor is <a href="/docs">Docs</a>, the value is /docs. Cheerio’s manipulation guide documents this attribute pattern at Manipulating the DOM.
Read other attributes from that same element in the usual way:
const first = $('a').first();
const href = first.attr('href');
const label = first.text().trim();
const title = first.attr('title');
attr('href') returns undefined when the selection is empty or the matched anchor has no href. Check both conditions explicitly when input is uncontrolled:
Free tools Windows power users keep installed
One-click scans. No signup required.
const link = $('a').first();
if (link.length === 0) {
console.log('No anchor found');
} else {
const href = link.attr('href');
console.log(href === undefined ? 'Anchor has no href' : href);
}
Get every href as a plain array
Map over the entire selection and finish with .get(). The final call converts Cheerio’s collection into a normal JavaScript array:
const hrefs = $('a')
.map((_, element) => $(element).attr('href'))
.get();
console.log(hrefs);
For anchors without an href, the mapped value is undefined. If your output should contain only actual values, filter them:
const hrefs = $('a')
.map((_, element) => $(element).attr('href'))
.get()
.filter((href) => typeof href === 'string' && href.length > 0);
Use a CSS selector when you need a narrower set, such as navigation links or links inside a card:
Rank #2
const navigationLinks = $('nav a')
.map((_, element) => ({
href: $(element).attr('href'),
text: $(element).text().trim(),
}))
.get();
The selector syntax and traversal patterns are covered in Selecting Elements and Traversing the DOM.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Choose between raw hrefs and absolute URLs
Raw attribute values with attr()
attr('href') returns exactly what appeared in the source. A relative value remains relative:
const $ = cheerio.load('<a href="../images/logo.svg">Logo</a>');
console.log($('a').attr('href')); // ../images/logo.svg
This method does not normalize, validate, fetch, or follow the URL. It is the right choice when you need to preserve the publisher’s original markup.
Resolved values with prop()
To resolve a relative link, supply a document URL and read the property:
const $ = cheerio.load('<a href="/docs">Docs</a>', {
baseURI: 'https://example.com/articles/page.html',
});
console.log($('a').prop('href')); // https://example.com/docs
Without a document URL, Cheerio has no origin against which to resolve /docs. When loading directly from a URL, fromURL provides that context automatically:
const $ = await cheerio.fromURL('https://example.com/articles/page.html');
const absoluteLinks = $('a')
.map((_, element) => $(element).prop('href'))
.get();
Use attr and prop deliberately: the former preserves the source string, while the latter is URL-aware when a base is available. The behavior is documented in Manipulating the DOM and Troubleshooting.
Extract links with Cheerio’s extract API
For a declarative shape, Cheerio’s extract method can collect all matching hrefs:
const data = $.extract({
links: [{ selector: 'a', value: 'href' }],
});
console.log(data.links); // ['/docs', '/blog']
An array descriptor means “collect every match.” A descriptor without the array returns the first match:
const data = $.extract({
firstLink: { selector: 'a', value: 'href' },
links: [{ selector: 'a', value: 'href' }],
});
The value: 'href' descriptor uses Cheerio’s property API, so URL resolution depends on whether the loaded document has a URL. For nested records and several fields, this approach keeps the output shape in one place. See Extracting Data with the extract Method.
Build a complete link extractor
The following function accepts HTML, optionally resolves URLs against a page URL, keeps link text, and removes anchors that have no href:
import * as cheerio from 'cheerio';
export function getLinks(html, pageUrl) {
const options = pageUrl ? { baseURI: pageUrl } : undefined;
const $ = cheerio.load(html, options);
return $('a')
.map((_, element) => {
const anchor = $(element);
const rawHref = anchor.attr('href');
if (!rawHref) return null;
return {
href: pageUrl ? anchor.prop('href') : rawHref,
rawHref,
text: anchor.text().replace(/s+/g, ' ').trim(),
};
})
.get()
.filter(Boolean);
}
const links = getLinks(
'<a href="/docs"> Docs </a><a href="mailto:help@example.com">Email</a>',
'https://example.com/guide/start.html',
);
console.log(links);
Keeping both rawHref and the resolved href is useful for audits: you can see what the page authored and what URL your crawler will use.
Load HTML from a file or an HTTP response
Local file
import { readFile } from 'node:fs/promises';
import * as cheerio from 'cheerio';
const html = await readFile('./page.html', 'utf8');
const $ = cheerio.load(html);
const links = $('a').map((_, el) => $(el).attr('href')).get();
console.log(links);
HTTP response
Fetch the page separately, then pass the returned HTML to Cheerio. Set baseURI to the final page URL when you want absolute links:
import * as cheerio from 'cheerio';
const response = await fetch('https://example.com/guide');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const html = await response.text();
const $ = cheerio.load(html, { baseURI: response.url });
const links = $('a').map((_, el) => $(el).prop('href')).get();
Network policy, redirects, authentication, retries, and response-size limits belong in the fetching layer. Cheerio only parses the string you pass to it.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallHandle fragments and document structure
cheerio.load treats input as a complete document by default and may add missing elements such as <html>, <head>, and <body>. For a fragment where that distinction matters, use the third argument:
Rank #4
const fragment = cheerio.load(
'<a href="/one">One</a><a href="/two">Two</a>',
null,
false,
);
const links = fragment('a').map((_, el) => fragment(el).attr('href')).get();
The fragment behavior and related parsing issues are described in Cheerio’s troubleshooting guide.
Why a link may be missing
The selector matched nothing
Check the count before reading attributes:
const anchors = $('a');
console.log('anchors:', anchors.length);
An empty selection makes attr() return undefined. Confirm that the supplied string contains the expected markup and that the selector is not limited to the wrong container.
The anchor has no href
Some sites use <a> as a JavaScript control or put a destination in data-href. Cheerio will not invent an href. Read the actual attribute you need and decide whether such elements should be excluded.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsYou expected an absolute URL
A result such as /docs is correct for attr('href'). Add baseURI (or load with fromURL) and use prop('href') when an absolute value is required.
The page creates links with JavaScript
Cheerio does not execute scripts. If the link appears only after hydration, scrolling, a click, or an API request in a browser, it will not be present in the static HTML. Use browser automation such as Puppeteer or Playwright to obtain rendered HTML, or use a DOM environment when appropriate, then pass that HTML to Cheerio.
The response is not the page you thought you fetched
Log the final response URL, status, content type, and a short prefix of the response body. Login pages, bot checks, consent pages, and error documents can contain few or no useful anchors even though parsing succeeded.
Performance, correctness, and operational notes
- Parse once and reuse the same
$function for all selectors on a document. - Use a narrow selector such as
main awhen navigation, footer, or hidden template links are irrelevant. - Use
.get()only when you need a plain array; keep the Cheerio collection while chaining filters and traversal operations. - For large documents, avoid retaining whole HTML strings and derived objects longer than necessary.
- Deduplicate only when your application wants unique destinations; repeated links can be meaningful in menus and content.
- Treat
mailto:,tel:, fragment-only values, and JavaScript URLs as different schemes rather than assuming every href is an HTTP page. - Resolve URLs before deduplication if
/docsandhttps://example.com/docsshould count as the same destination.
Or skip the browser setup
If your goal is to obtain clean page HTML or screenshots before processing links, ScreenshotNeo can fetch a URL through one API request. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Use the URL directly (adapt the target URL as needed):
Best Value
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
'https://api.screenshotneo.com/v1/shot',
params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', buffer));
See the ScreenshotNeo API documentation for options such as full-page capture, waiting for a selector or network idle, custom headers and cookies, blocking resources, geolocation, caching, asynchronous jobs, bulk capture, and PDF output. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to get an API key.
FAQ
Does Cheerio crawl every link automatically?
No. It extracts strings from markup. Your application must decide which URLs to request, enforce robots and rate-limit policies, and handle redirects or failures.
Can I extract links from XML with Cheerio?
Cheerio is primarily used for HTML, but its parser options can be configured for XML-style input. Test the exact document and selectors you receive because XML is case-sensitive and may use namespaces.
Recommended Free Tools
How do I preserve duplicate links?
Do not put the results in a Set or deduplicate them. The order returned by map follows document order, so repeated navigation and content links remain visible.
Frequently Asked Questions
Does Cheerio crawl every link automatically?
No. Cheerio extracts href strings from the markup you provide; your program must request URLs and handle robots rules, rate limits, redirects, and errors.
Can I extract links from XML with Cheerio?
Cheerio is primarily used for HTML. XML-style parsing can be configured, but XML case sensitivity and namespaces require selectors suited to the document you receive.
How do I preserve duplicate links?
Keep the array returned by map without converting it to a Set or otherwise deduplicating it; values remain in document order.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

