Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The dependable starter workflow is simple: use n8n’s HTTP Request node to download a page, then pass the returned HTML to the HTML node and extract fields with CSS selectors. This works when the server response already contains the content you need. It is not, by itself, proof that n8n renders JavaScript-generated content like a full browser.

Before you scrape

Confirm permission and scope

Check the target site’s terms, robots guidance where relevant, applicable law, and any contractual restrictions before collecting or reusing data. A successful HTTP response only proves that the server answered; it does not grant permission to copy, store, or redistribute the content. Prefer an official API when one provides the fields you need.

Inspect one real response

Open the target URL with a normal HTTP client or browser and determine whether the desired text is present in the initial HTML response. If the browser shows a list that is absent from “view source” or the HTTP response body, the site may be filling it with JavaScript after load. Treat that as a tool-selection question rather than assuming the basic n8n pattern will work.

Build the basic n8n scraper

1. Create a workflow and add HTTP Request

  1. Create a new workflow and add an HTTP Request node.
  2. Set the method to GET.
  3. Enter the page URL, such as https://example.com/products.
  4. Choose a response format that preserves the body (normally text). Enable status and headers when you are diagnosing failures.
  5. Add authentication, query parameters, custom headers, cookies, a proxy, a timeout, or redirect handling only when the target requires them. Keep credentials in n8n’s credential system or environment variables instead of hard-coding secrets in expressions.
  6. Execute the node and inspect the result. Verify that the response is the expected page, not a login screen, consent wall, error document, or challenge.

The node supports request methods, authentication, headers, batching, pagination, proxy settings, and timeout controls. Use the smallest request configuration that the target permits; extra headers or aggressive concurrency can cause blocking.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Add the HTML node

  1. Connect an HTML node to HTTP Request. In n8n 0.213.0 and later, HTML is the current node name; older tutorials may call the node HTML Extract.
  2. Set the input property to the HTTP Request field containing the response body (often body; confirm by viewing the node output).
  3. Add an extraction operation and enter a CSS selector, for example article h2, .price, or main a.product-link.
  4. Select the output type: text, inner HTML, an attribute such as href, or a form value.
  5. Enable array output when a selector can match several elements. Trim or clean whitespace after extraction if the target includes line breaks or navigation text.

A useful item might contain title, url, and price fields, each extracted with a separate selector. Test selectors against the actual response rather than guessing from a screenshot.

3. Normalize and validate records

Use a Code node only for transformation: trimming strings, converting prices, joining relative URLs to a base URL, dropping duplicates, or adding validation flags. n8n documents that Code is not the node for making HTTP requests; use HTTP Request for network access. A defensive JavaScript transformation can mark missing fields instead of silently emitting incomplete records:

return items.map(item => {
  const d = item.json;
  return {
    json: {
      title: typeof d.title === 'string' ? d.title.trim() : null,
      url: typeof d.url === 'string' ? d.url.trim() : null,
      missing: ['title', 'url'].filter(k => !d[k])
    }
  };
});

Route records with missing required fields to an error branch or review queue. Do not treat an HTTP 200 response as a valid scrape until the body and required selectors have been checked.

Pagination, batching, and rate control

Page-number and next-link pagination

Inspect one response first. Identify whether the target uses a page query such as ?page=2, an offset, a cursor, or a next-page URL. Configure HTTP Request pagination to update the relevant parameter or follow the returned next URL. Pagination behavior and limits are target-dependent, so copy the target’s documented rules rather than assuming page numbers.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Stop when the response has no next link, the API returns an explicit end condition, or a safe maximum is reached. Record the page or cursor with each item so a failed run can resume without duplicating earlier data.

Independent URLs

For a list of unrelated URLs, feed one URL per item into HTTP Request and use batching or an interval between requests. Start conservatively, watch for 429 responses, and honor the service’s published limits. A retry should have a bounded count and increasing delay; otherwise a temporary failure can become an accidental denial-of-service pattern.

API pagination versus HTML pagination

An official API is usually preferable when it exposes the required fields with stable identifiers and documented pagination. HTML scraping is appropriate when the information is public, permitted, and unavailable through an API, but selectors can break whenever the site redesigns its markup.

Approach Best fit Watch for
Official API Structured data, authentication, documented limits Quota, fields not exposed, version changes
HTTP Request + HTML Server-rendered pages with stable selectors Markup changes, consent pages, rate limits
Browser-capable capture Pages that require execution before content appears Higher complexity, browser failures, bot checks

JavaScript-rendered pages: know the boundary

The HTTP Request plus HTML combination parses the response it receives. The reviewed n8n documentation describes request and extraction controls, but does not establish that this basic pair launches a browser and renders JavaScript-generated content. If the required elements are missing from the response body, confirm whether the site has an API or server-rendered endpoint. If browser execution is necessary, choose a browser-capable service or automation tool and pass its resulting HTML or data into n8n.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not “fix” missing content by adding arbitrary delays to HTTP Request; a delay does not execute the page’s JavaScript. Similarly, a Code node cannot make the network call that the HTTP Request node should make.

Code-node and hosting considerations

n8n’s Code documentation distinguishes Cloud and self-hosted environments. Package imports and external modules may be available on self-hosted installations when explicitly enabled, while Cloud has restrictions. The documentation describes Pyodide as a legacy Python option and documents native Python support in newer releases. Because these capabilities vary by n8n version and hosting, check the documentation matching your installation before adopting a Python-based example. For a portable workflow, keep HTTP access in HTTP Request and use built-in Code functionality for small transformations.

Make the workflow fail visibly

  • Check the HTTP status and content type before extraction.
  • Reject bodies that contain a login page, CAPTCHA, bot challenge, or generic error template.
  • Validate that each required selector returned at least one value.
  • Store the source URL, retrieval time, page or cursor, and parser version with each record.
  • Send failures to a separate branch with the response status and a short body sample for diagnosis.
  • Re-test selectors after target redesigns; do not silently accept empty arrays.

Troubleshooting common failures

HTTP 401 or 403

The target requires authentication or rejects the request. Use the permitted authentication method, required headers, or cookies; do not attempt to bypass an access control. Confirm that your use complies with the site’s rules.

HTTP 429

You are being rate-limited. Reduce concurrency, add an interval, honor the published limit, and implement bounded retries with backoff. A proxy is not a justification for ignoring limits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HTTP 200 but extracted fields are empty

Inspect the raw body. The selector may be wrong, the content may be JavaScript-generated, or you may have received a consent, login, or challenge page. Test a selector against the exact response stored by n8n.

Only the first match is returned

Enable array output in the HTML node for selectors that match multiple elements, then iterate over the resulting values or split them into items.

Redirect or timeout errors

Inspect redirect settings and the final URL. Increase the timeout only when the target legitimately responds slowly; otherwise fix the URL, reduce payload size, or use the target’s API. Record failures instead of converting them to empty records.

Relative links or unwanted whitespace

Extract the href attribute, resolve it against the page’s base URL in a Code node, and trim or normalize text before storage.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Operational checklist

  1. Permission and data-use rules checked.
  2. One response inspected for the required fields.
  3. HTTP Request configured with the smallest necessary authentication and headers.
  4. HTML selectors tested against real markup, with array output where needed.
  5. Pagination stop conditions and request pacing defined.
  6. Missing fields, non-success responses, and challenge pages routed to errors.
  7. Run history, source URL, cursor, and retrieval time retained for auditing.
  8. Selectors monitored after site changes.

Or skip the browser setup

If you need a clean screenshot or PDF as part of an n8n workflow, ScreenshotNeo provides a single HTTP endpoint and an MCP server for AI agents. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing result in headers. It is not a replacement for extracting structured fields with CSS selectors, but it can supply a rendered visual artifact when a plain HTTP response is insufficient.

Use the API details in the ScreenshotNeo documentation. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page capture, lazy-image loading, CSS-selector element capture, device and viewport controls, retina scale, PDF settings, custom CSS and JavaScript, waits, request blocking, cookies and headers, geolocation, transparent backgrounds, caching, signed links, async webhooks, bulk capture, usage data, and an OpenAPI specification. Its free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Sign up free for ScreenshotNeo.

FAQ

Does n8n scrape a website automatically?

No. You build the request, extraction, pagination, validation, and storage steps, and you remain responsible for permission and maintenance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which n8n node replaced HTML Extract?

The HTML node replaced HTML Extract in n8n 0.213.0. Older tutorials may therefore show a different node name.

Can the Code node fetch a URL?

Use HTTP Request for HTTP access. Use Code for transformations and logic after the response arrives.

Should I use Cloud or self-hosted n8n?

Choose based on required package imports, Python behavior, network access, credential controls, and operational ownership. The available Code features differ between hosting models and versions, so verify the documentation for your installation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.