Custom rules turn a browser API into a scraper by describing the page-specific actions needed to reveal data. The remote browser loads the site, runs those actions—such as typing, clicking, scrolling, waiting, or executing JavaScript—and returns the resulting HTML or structured fields for parsing. The API supplies the execution environment; your rules supply the workflow.
What “custom rules” add to a browser API
A conventional HTTP scraper requests a URL and parses the response it receives. That is enough when the data is present in the initial HTML. Many modern sites instead build the useful page state in a browser: JavaScript makes additional requests, inserts results into the DOM, and reveals content only after a user action.
Custom rules describe those actions for one target site. Oxylabs’ description of “Custom Browser Instructions” follows this model: submit website-specific instructions, let the browser execute them, then receive the result as raw HTML or structured JSON. The exact syntax differs by provider, but the division of responsibility is consistent.
- Browser API: supplies a remote, script-capable browser session.
- Rules: specify navigation, controls, waits, and extraction conditions for the target page.
- Your parser or downstream system: checks the returned result and stores the fields you need.
The inspect–interact–wait–extract workflow
1. Inspect the target page
Open the real page and identify both the data and the controls that expose it. Note stable selectors for search fields, submit buttons, tabs, dropdowns, result cards, pagination, and any element that signals completion. Inspect the DOM after the page has finished rendering, not only the source returned by an initial request.
#1 Best Overall
2. Write the interaction sequence
Translate the observed journey into ordered rules. A typical sequence might fill a search field, click submit, scroll to trigger lazy loading, and wait for a result selector. Providers may also support JavaScript execution, custom waits, request interception, or navigation commands. Keep each action explicit so an individual failure can be diagnosed.
3. Wait for the page state you actually need
A fixed delay is a weak substitute for a condition. When supported, wait for the target element to appear, for a relevant request to finish, or for network activity to become idle. Dynamic elements can arrive at different speeds, so extracting immediately after a click can return an empty or partial page.
Rank #2
4. Extract and validate
Retrieve the rendered HTML or the service’s structured response, then verify that expected fields exist and contain plausible values. Treat an HTTP success response as transport success, not proof that the scraper found the right data. Record missing selectors, empty result sets, and unexpected layouts for review.
5. Test against the real target before production
Run the rules repeatedly against the target site and representative inputs. Web Scraper’s documentation cautions that no universal tool guarantees compatibility with every website. Selector validation, retries, and monitoring are operating requirements, not optional polish.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →When browser automation is worth the overhead
Use a browser when the page requires state or interaction
- Results appear only after JavaScript executes.
- A click, form fill, dropdown selection, or scroll reveals the data.
- The site is a single-page application that updates without a full navigation.
- You need to observe page XHR or fetch requests made after an interaction.
- You already have Puppeteer, Playwright, or Selenium logic and want a managed remote browser.
Prefer a lighter HTTP method for simple pages
If the required fields are present in the initial response and no interaction is needed, a browser adds startup time, resource use, and another layer of failure. Bright Data’s reference separates simple HTTP scraping from its Browser API use cases such as clicking, scrolling, filling forms, running JavaScript, handling single-page applications, and intercepting XHR or fetch requests. That is vendor guidance rather than a universal performance benchmark.
Common rule failures and practical checks
Selector mismatch
A renamed class, changed nesting, or different result template can make an action miss its target. Scrape.do describes per-action success and error reporting; use equivalent diagnostics where available and fail the job clearly instead of silently storing an empty result.
Extraction starts too soon
If a fixed delay ends before the request or rendering completes, the scraper captures an incomplete state. Prefer a selector- or request-based wait tied to the data you need.
The page changed
Sites routinely alter controls and layouts. Revalidate rules against the target, monitor field presence, and keep selectors as specific as necessary but no more brittle than required.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Best Value
Browser mode changes behavior
Desktop and mobile sessions can expose different controls. Scrape.do notes that its Android-based mobile browser infrastructure uses a Tap action because Click does not work there. Treat device mode as part of the rule design and test each mode you intend to run.
Anti-bot or consent interstitials
A challenge, consent screen, or login wall can replace the expected page. Detect these states explicitly and follow the target site’s terms and access requirements; do not assume that a rule written for the normal page also handles an interstitial.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choosing an implementation approach
| Approach | How it works | What to compare |
|---|---|---|
| Custom-instruction scraping API | Submit site-specific browser actions; the provider renders the page and returns HTML or structured JSON. | Supported actions, output format, wait behavior, maintenance, and current service price. |
| Framework-connected cloud browser | Connect Puppeteer, Playwright, or Selenium to a managed browser session. | Framework support, session setup, debugging access, control, and operational complexity. |
| Sitemap-based extension or cloud service | Define navigation and selectors in a sitemap; hosted features may add scheduling and delivery. | Local versus hosted execution, selector validation, scheduling, retries, and export. |
| Trained-agent scraper | Train an agent to capture named fields and invoke it through an API, webhook, or polling workflow. | Setup effort, field structure, adaptation to page changes, and integration options. |
These categories solve different problems. Compare them on the same target pages, required output, interaction sequence, maintenance expectations, and current plan terms rather than assuming one category is universally best.
A maintainable rule design
- Define the output contract: list every field, its expected type, and what counts as missing.
- Separate navigation from extraction: keep steps that reach the state distinct from selectors that read values.
- Use meaningful waits: tie completion to a result element or request whenever possible.
- Capture action-level diagnostics: retain which step failed and the page state returned.
- Test multiple states: include empty searches, pagination, slow responses, mobile layouts, and pages with no matching records.
- Monitor for drift: alert when required fields disappear or their formats change.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, not a general-purpose structured-data scraper. When your workflow needs a clean visual capture of the rendered result, one GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server lets Claude, Cursor, or another MCP client use take_screenshot, get_page_info, and capture_pdf.
Free tools Windows power users keep installed
One-click scans. No signup required.
Example using the documented endpoint (see the ScreenshotNeo API documentation):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is included on every plan. Sign up free.

