Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Custom rules turn a browser API into a scraper by describing the page-specific actions needed to reveal data. The remote browser loads the site, runs those actions—such as typing, clicking, scrolling, waiting, or executing JavaScript—and returns the resulting HTML or structured fields for parsing. The API supplies the execution environment; your rules supply the workflow.

What “custom rules” add to a browser API

A conventional HTTP scraper requests a URL and parses the response it receives. That is enough when the data is present in the initial HTML. Many modern sites instead build the useful page state in a browser: JavaScript makes additional requests, inserts results into the DOM, and reveals content only after a user action.

Custom rules describe those actions for one target site. Oxylabs’ description of “Custom Browser Instructions” follows this model: submit website-specific instructions, let the browser execute them, then receive the result as raw HTML or structured JSON. The exact syntax differs by provider, but the division of responsibility is consistent.

  • Browser API: supplies a remote, script-capable browser session.
  • Rules: specify navigation, controls, waits, and extraction conditions for the target page.
  • Your parser or downstream system: checks the returned result and stores the fields you need.

The inspect–interact–wait–extract workflow

1. Inspect the target page

Open the real page and identify both the data and the controls that expose it. Note stable selectors for search fields, submit buttons, tabs, dropdowns, result cards, pagination, and any element that signals completion. Inspect the DOM after the page has finished rendering, not only the source returned by an initial request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Write the interaction sequence

Translate the observed journey into ordered rules. A typical sequence might fill a search field, click submit, scroll to trigger lazy loading, and wait for a result selector. Providers may also support JavaScript execution, custom waits, request interception, or navigation commands. Keep each action explicit so an individual failure can be diagnosed.

3. Wait for the page state you actually need

A fixed delay is a weak substitute for a condition. When supported, wait for the target element to appear, for a relevant request to finish, or for network activity to become idle. Dynamic elements can arrive at different speeds, so extracting immediately after a click can return an empty or partial page.

4. Extract and validate

Retrieve the rendered HTML or the service’s structured response, then verify that expected fields exist and contain plausible values. Treat an HTTP success response as transport success, not proof that the scraper found the right data. Record missing selectors, empty result sets, and unexpected layouts for review.

5. Test against the real target before production

Run the rules repeatedly against the target site and representative inputs. Web Scraper’s documentation cautions that no universal tool guarantees compatibility with every website. Selector validation, retries, and monitoring are operating requirements, not optional polish.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When browser automation is worth the overhead

Use a browser when the page requires state or interaction

  • Results appear only after JavaScript executes.
  • A click, form fill, dropdown selection, or scroll reveals the data.
  • The site is a single-page application that updates without a full navigation.
  • You need to observe page XHR or fetch requests made after an interaction.
  • You already have Puppeteer, Playwright, or Selenium logic and want a managed remote browser.

Prefer a lighter HTTP method for simple pages

If the required fields are present in the initial response and no interaction is needed, a browser adds startup time, resource use, and another layer of failure. Bright Data’s reference separates simple HTTP scraping from its Browser API use cases such as clicking, scrolling, filling forms, running JavaScript, handling single-page applications, and intercepting XHR or fetch requests. That is vendor guidance rather than a universal performance benchmark.

Common rule failures and practical checks

Selector mismatch

A renamed class, changed nesting, or different result template can make an action miss its target. Scrape.do describes per-action success and error reporting; use equivalent diagnostics where available and fail the job clearly instead of silently storing an empty result.

Extraction starts too soon

If a fixed delay ends before the request or rendering completes, the scraper captures an incomplete state. Prefer a selector- or request-based wait tied to the data you need.

The page changed

Sites routinely alter controls and layouts. Revalidate rules against the target, monitor field presence, and keep selectors as specific as necessary but no more brittle than required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Browser mode changes behavior

Desktop and mobile sessions can expose different controls. Scrape.do notes that its Android-based mobile browser infrastructure uses a Tap action because Click does not work there. Treat device mode as part of the rule design and test each mode you intend to run.

Anti-bot or consent interstitials

A challenge, consent screen, or login wall can replace the expected page. Detect these states explicitly and follow the target site’s terms and access requirements; do not assume that a rule written for the normal page also handles an interstitial.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choosing an implementation approach

Approach How it works What to compare
Custom-instruction scraping API Submit site-specific browser actions; the provider renders the page and returns HTML or structured JSON. Supported actions, output format, wait behavior, maintenance, and current service price.
Framework-connected cloud browser Connect Puppeteer, Playwright, or Selenium to a managed browser session. Framework support, session setup, debugging access, control, and operational complexity.
Sitemap-based extension or cloud service Define navigation and selectors in a sitemap; hosted features may add scheduling and delivery. Local versus hosted execution, selector validation, scheduling, retries, and export.
Trained-agent scraper Train an agent to capture named fields and invoke it through an API, webhook, or polling workflow. Setup effort, field structure, adaptation to page changes, and integration options.

These categories solve different problems. Compare them on the same target pages, required output, interaction sequence, maintenance expectations, and current plan terms rather than assuming one category is universally best.

A maintainable rule design

  1. Define the output contract: list every field, its expected type, and what counts as missing.
  2. Separate navigation from extraction: keep steps that reach the state distinct from selectors that read values.
  3. Use meaningful waits: tie completion to a result element or request whenever possible.
  4. Capture action-level diagnostics: retain which step failed and the page state returned.
  5. Test multiple states: include empty searches, pagination, slow responses, mobile layouts, and pages with no matching records.
  6. Monitor for drift: alert when required fields disappear or their formats change.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not a general-purpose structured-data scraper. When your workflow needs a clean visual capture of the rendered result, one GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server lets Claude, Cursor, or another MCP client use take_screenshot, get_page_info, and capture_pdf.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Example using the documented endpoint (see the ScreenshotNeo API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is included on every plan. Sign up free.