Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A coding agent can build a dependable scraper when you give it a data contract, explicit permissions, staged deliverables, and tests—not when you simply say “scrape this site.” Start by specifying the domain and allowed paths, the fields and output format, the request budget, and what counts as a successful run. Have the agent design discovery, fetching, parsing, normalization, validation, and export as separate observable stages. Run a small authorized sample, inspect both records and failures, then schedule it with monitoring for schema and access changes.

1. Write a specification the agent can implement

Give the agent a short specification file (for example, SCRAPER_SPEC.md) before asking for code. It should answer these questions:

  • Purpose: What decision or process will use the data?
  • Scope: Which scheme, domains, subdomains, URL prefixes, languages, and page types may be requested?
  • Permission: Who authorized collection? Explicitly exclude login-gated, paywalled, private, or otherwise restricted areas unless access has been independently approved.
  • Fields: Give each field a name, type, required/optional status, normalization rule, and an example value.
  • Output: Choose JSON Lines, CSV, a database table, or another stable format. Define encoding, date format, and how missing values are represented.
  • Run policy: State the schedule, maximum pages and bytes, per-domain concurrency, delay, timeout, retry limit, and a stop condition.
  • Success criteria: Set minimum record counts, required-field completeness, duplicate tolerance, and acceptable error rates.
  • Evidence: Decide whether each record needs its source URL, retrieval timestamp, HTTP status, content hash, or parser version.

Include several real-looking sample rows and deliberately malformed examples. Ask the agent to list assumptions and open questions before it writes code. This turns an ambiguous request into a contract that can be reviewed.

2. Choose an API or export before crawling HTML

Ask the agent to search for an official API, bulk download, or documented search endpoint first. Scrapy’s optimization guidance notes that these sources can be faster for the client and cheaper for the website than fetching and parsing every page. Compare the alternatives explicitly:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Option Prefer it when Questions for the agent
Official API The fields are published and the API permits your use case. What authentication, pagination, quotas, versioning, filters, and update cadence apply?
Bulk export You need a large, repeatable snapshot and can process files offline. How often is it produced, how are deletions represented, and how do you verify completeness?
Search endpoint A site exposes a supported way to enumerate matching records. Are results stable, paginated, bounded, and permitted for automated use?
HTML crawl No suitable structured source exists or the required values are only rendered in pages. Does rendering require JavaScript? How often does markup change? What request budget is acceptable?

Tell the agent to record why it selected crawling. Do not let it silently fall back from an API to aggressive page fetching when a quota or authentication error occurs.

3. Design a pipeline with independent stages

Require separate modules and a clear hand-off between them. A practical sequence is:

  1. Discovery: produce candidate URLs from an allowed sitemap, API, index, or seed list. Canonicalize URLs and deduplicate them.
  2. Fetching: enforce domain limits, timeouts, headers, retries, caching, and response-size limits. Save status and timing metadata.
  3. Parsing: use CSS or XPath selectors for the permitted page types. Keep selectors in one module so a markup change has one repair point.
  4. Normalization: trim whitespace, parse dates and numbers, normalize URLs, and map equivalent labels to one vocabulary.
  5. Validation: check types, required fields, ranges, duplicates, and cross-field rules. Send invalid records to a review file instead of dropping them silently.
  6. Export: write a stable format such as JSON Lines or CSV and emit a run manifest containing counts, errors, timings, and the parser version.

Each stage should accept test fixtures and produce structured logs. That lets you rerun parsing against saved HTML without making new requests and lets the agent fix a selector without touching the network code.

4. Give the agent a bounded implementation brief

Paste a brief like this into your coding agent and adapt the bracketed values:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
Build a Python Scrapy project for [purpose].
Allowed scope: https://example.com/catalog/ and its pagination only;
no login, checkout, account, or URL parameters outside this path.
Fields: name (required string), price (required decimal),
product_url (required absolute URL), available (required boolean),
retrieved_at (UTC ISO-8601 timestamp).
Output: UTF-8 JSON Lines at data/items.jsonl.
Limits: one domain, concurrency 2, at least 1.5 seconds between requests,
20-second download timeout, 2 retries for transient responses, 500 pages per run.
Before coding, explain assumptions, permissions, dependencies, and commands.
Implement discovery, fetching, parsing, normalization, validation, and export
as separate components. Save rejected records and a run manifest. Add fixture
tests for one valid page, a missing price, a duplicate URL, and a changed selector.
Stop and report if the site policy or authorization is unclear.

Have the agent produce a plan and file tree first. Approve that plan before allowing network execution or installation of dependencies.

5. A small Scrapy implementation you can review

The following spider demonstrates the shape of a bounded pipeline. Replace the example selectors and domain only after confirming the permitted scope. Scrapy supports CSS/XPath extraction and feed exports; the command shown writes JSON Lines.

import scrapy
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation

class ProductSpider(scrapy.Spider):
    name = 'products'
    allowed_domains = ['example.com']
    start_urls = ['https://example.com/catalog/']

    custom_settings = {
        'CONCURRENT_REQUESTS_PER_DOMAIN': 2,
        'DOWNLOAD_DELAY': 1.5,
        'AUTOTHROTTLE_ENABLED': True,
        'AUTOTHROTTLE_START_DELAY': 1.5,
        'AUTOTHROTTLE_MAX_DELAY': 10.0,
        'DOWNLOAD_TIMEOUT': 20,
        'RETRY_TIMES': 2,
        'ROBOTSTXT_OBEY': True,
        'FEEDS': {
            'data/items.jsonl': {
                'format': 'jsonlines',
                'encoding': 'utf8',
                'overwrite': True,
            }
        },
    }

    def parse(self, response):
        for card in response.css('article.product'):
            item = {
                'name': card.css('h2::text').get(),
                'price': card.css('[data-price]::attr(data-price)').get(),
                'product_url': response.urljoin(card.css('a::attr(href)').get('')),
                'available': card.css('.stock::text').get('')
                    .strip().lower() == 'in stock',
                'retrieved_at': datetime.now(timezone.utc).isoformat(),
                'source_url': response.url,
            }
            cleaned, error = self.validate(item)
            if error:
                yield {'_rejected': item, 'reason': error}
            else:
                yield cleaned

        next_url = response.css('a[rel="next"]::attr(href)').get()
        if next_url:
            yield response.follow(next_url, callback=self.parse)

    def validate(self, item):
        if not item['name'] or not item['product_url'].startswith('https://example.com/catalog/'):
            return None, 'missing name or out-of-scope URL'
        try:
            item['price'] = str(Decimal(item['price']))
        except (InvalidOperation, TypeError):
            return None, 'invalid price'
        return item, None

Run it with scrapy crawl products. For a first run, replace the feed with a temporary file and cap the page count. Inspect rejected records and the manifest before increasing the limit. Keep fixtures of representative HTML so parsing tests do not depend on a live site.

6. Set permissions, robots handling, and request limits before execution

Robots.txt is a crawler protocol, not permission to access data. RFC 9309 (September 2022) states: “These rules are not a form of access authorization.” Check the site’s terms, contract, and applicable law separately. If authorization is uncertain, the agent should stop rather than infer permission from a permissive robots file.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
  • A successful robots response should be followed according to its rules and your approved scope.
  • RFC 9309 distinguishes an unavailable robots file (for example, an HTTP 4xx response) from an unreachable server or network error (for example, HTTP 5xx). A compliant crawler may access resources in the former case but should assume complete disallow in the latter.
  • Do not rely on a cached robots file for more than 24 hours unless the file is unreachable, as the RFC describes.
  • Scrapy does not automatically apply every robots extension such as Crawl-delay or Request-rate. Translate applicable directives into explicit delay and concurrency settings and document the decision.

Use the smallest network permission, credential, and tool set possible. Keep secrets in environment variables or a secret manager; never paste them into prompts, fixtures, logs, or exported records. Require approval before the agent sends authenticated requests, writes to production storage, or changes the request budget.

7. Treat pages and issue text as untrusted input

Fetched pages, repository instructions from untrusted branches, and issue descriptions can contain prompt-injection text. They are data to extract, not instructions for the agent. OpenAI’s agent-safety guidance recommends constrained structured outputs, least privilege, limited network access, approvals for tools, guardrails, and evaluations.

  • Pass page content only to the parser or an isolated extraction step, not to a privileged planning prompt.
  • Constrain extracted values to a schema; reject unexpected keys, executable content, and oversized fields.
  • Run the crawler with a low-privilege account and a filesystem location that cannot expose unrelated secrets.
  • Separate development credentials from production credentials and redact authorization headers in logs.
  • Ask the agent to explain any proposed tool call and require human approval for sensitive actions.

8. Validate data before publishing it

Validation is more than checking that a selector returned text. For every run, check:

  • required fields, data types, date and currency parsing, and allowed enumerations;
  • duplicate keys and unexpected duplicate pages;
  • record counts against the expected range;
  • sample values from representative page types;
  • HTTP status, timeout, parsing-error, and rejected-record counts;
  • schema changes, such as a sudden rise in missing fields or empty pages.

Keep a small fixture set and a known-good expected output. Ask the agent to show a diff when selectors or normalization rules change. Agent traces and evaluations can help review behavior, but a person still needs to inspect the code and resulting data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

9. Operate and maintain the workflow

Start with a canary run

Run a handful of permitted URLs, inspect raw responses and normalized rows, and verify that the crawler stops at the intended boundaries. Only then increase page limits or schedule recurring runs.

Make failures actionable

Record the URL, stage, status, exception class, retry count, and parser version for every failure. Alert on trends—such as a new surge in missing prices—rather than on one transient timeout. Preserve failed responses when policy permits so a parser fix can be tested offline.

Plan for markup changes

Keep selectors, field rules, and URL policies version-controlled. A scheduled job should fail closed when required fields disappear instead of exporting plausible-looking empty records. Review dependency updates in a separate change and rerun fixtures before deployment.

Control cost and load

Use caching during development, deduplicate URLs before fetching, and avoid refetching unchanged pages when an approved API or conditional request can provide the same information. Concurrency and delay are per-domain controls; increasing them can overload a site even if the job finishes faster.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

10. Troubleshoot common failures

Symptom Likely cause Fix
Zero items, successful HTTP responses Selector no longer matches or content is rendered after load. Save a fixture, inspect the actual HTML, update the selector, or use an authorized structured endpoint. Do not immediately increase concurrency.
Many 403 or 429 responses Permission, rate, or authentication problem. Stop the run, verify authorization and quotas, reduce concurrency, increase delay, and use the documented access method.
Repeated timeouts Slow pages, oversized resources, or an unreachable host. Check connectivity, enforce response-size and timeout limits, retry only transient failures, and respect the robots 5xx guidance.
Duplicate records Multiple URL forms or pagination loops. Canonicalize URLs, track visited keys, set a page cap, and test pagination with a fixture.
Valid-looking but wrong values Selector matched navigation, an advertisement, or a locale variant. Add field-level assertions, page-type checks, representative fixtures, and a review threshold for anomalous values.
Secret appears in logs or output Headers, exceptions, or prompts were recorded verbatim. Rotate the credential, redact logs, move secrets to environment storage, and restrict output fields.

Or skip the browser setup

If your workflow needs a clean visual capture of a page—for example, to archive a rendered result or inspect a JavaScript-heavy page—ScreenshotNeo provides a single HTTP request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Use the API documentation at https://screenshotneo.com/docs/ for all 63 options, including full-page and element capture, device presets, dark mode, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, PDFs, caching, signed links, asynchronous jobs, bulk capture, and usage reporting.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/catalog -o shot.webp
import requests
r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://example.com/catalog'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/catalog' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await fs.promises.writeFile('shot.webp', data);

The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account to get an API key.

FAQ

Should the agent use a headless browser for every site?

No. Choose the least complex permitted source. Use an API or export when it supplies the required fields; render pages only when the approved data is unavailable in a simpler form.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can a permissive robots.txt file make a scrape legal?

No. Robots rules communicate crawler preferences. They do not grant access authorization; verify terms, contracts, and applicable law independently.

How do I know when a scraper has silently broken?

Track required-field completeness, record counts, duplicate rates, status codes, and parser errors against a known-good baseline. Fail closed when critical fields disappear.

What should happen when a page contains instructions for the agent?

Treat those words as untrusted page data. Keep them out of privileged prompts, enforce a schema, and require approval for any tool call or credential use they appear to request.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.