To modify a web scrape with an API, change both sides of the data pipeline: send the right request to the API, then adapt your scraper to the response it actually returns. That means checking the endpoint’s authentication and input rules, mapping its JSON or HTML fields, handling pagination and errors, and validating the transformed records before saving them. If the page relies on JavaScript rather than a data API, use a documented rendering service or a browser-based approach instead of assuming an API request will reproduce the page.
What it means to modify a web scrape with an API
The phrase can describe two different changes. You may be replacing a scraper’s direct requests to a website with requests to a data API, or you may be sending a target URL to a hosted scraping API that fetches and returns page content for you. In either case, changing only the URL is not enough: the request format, authentication, response parser, pagination, and error handling must match the API contract.
A data API commonly returns structured JSON records. A rendered-page API may return HTML or a rendered screenshot, while a hosted scraper platform may run a job and provide its output later as a dataset. Those are different products and workflows. Before editing code, identify which kind of endpoint you have and what it promises to return.
- Data API: Request records directly from the service that owns the data. You usually parse JSON and implement pagination.
- Rendered-page API: Submit a web-page URL and options such as JavaScript rendering, headers, or proxy location. Parse the resulting HTML or other page output.
- Hosted scraper platform: Select a scraper or tool, start a run, check its status, then export its dataset. The platform can reduce browser and infrastructure work, but ties the integration to its job model, schema, limits, and pricing.
Scrapy.io describes the hosted pattern as a run followed by polling and dataset retrieval. WebScraping.AI documents rendered-page options such as JavaScript execution and custom headers; ScraperAPI documents rendering and proxy options. The exact inputs and response format depend on the particular service and endpoint.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Plan the request and response changes
Start with the API’s current documentation, not assumptions carried over from the old scraper. Record the method, endpoint, authentication scheme, required parameters, request body, response format, pagination signals, rate limits, and whether the result arrives synchronously or as a job. Do not reuse CSS selectors against a JSON response or send browser-only settings to an endpoint that does not document them.
- Identify the endpoint. Decide whether the target exposes a data API, or whether you need a service that fetches and renders the page. If the provider has a catalog endpoint, inspect it for the supported scraper or tool before creating a run.
- Move credentials out of source code. Use the authentication method the service documents, commonly an authorization bearer token or an API-key header. Store the secret in a server-side environment variable or secret manager; do not embed it in browser JavaScript, a public repository, or a URL that may be logged.
- Change inputs deliberately. Add only documented query parameters, headers, cookies, POST data, rendering flags, session settings, proxy, or country options. Confirm whether values belong in the query string, headers, or JSON body.
- Inspect a real response. Save a representative response and find the records array, nested fields, status or error object, and any continuation URL, cursor, total, offset, or limit. Check response headers too: some APIs communicate continuation or rate-limit information there.
- Define the output schema. Map source fields to names and types your application expects. Decide how to handle missing, malformed, duplicate, or newly added fields rather than allowing them to fail silently.
Keep the source contract separate from your application schema
Put the API-specific parsing in one function and translate the result into a stable internal record. For example, an API may return a price as a string and identify an item with product_id, while the rest of your application expects a numeric price and an id field. Normalize those differences at the boundary. When the provider changes its response, you can update one adapter instead of rewriting every consumer of the data.
Record a stable item key for deduplication. Log the source URL or endpoint, run or request identifier when available, and a useful error category. Avoid logging authorization headers, cookies, or other secrets. Save raw response samples securely when they are needed for debugging, and give them an appropriate retention period.
Example: paginate through a JSON API in Python
This Python example demonstrates the request-and-transform pattern for an API whose documented response contains items, total, offset, and limit. Scrapy.io documents this style of offset-and-limit pagination. Replace the endpoint and field mapping with the values in your own API’s documentation; the example is not a universal API contract. It uses a bearer token from the environment, enforces a timeout, raises on HTTP errors, stops on an empty page, and avoids requesting beyond the reported total.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
import os
import requests
API_URL = os.environ["API_URL"]
API_TOKEN = os.environ["API_TOKEN"]
PAGE_SIZE = 100
def normalize(item):
"""Adapt these fields and conversions to the documented response schema."""
if not isinstance(item, dict) or item.get("id") is None:
raise ValueError("Record is missing its required id")
return {
"id": str(item["id"]),
"name": str(item.get("name", "")),
}
def fetch_all():
headers = {"Authorization": f"Bearer {API_TOKEN}"}
offset = 0
results = []
while True:
response = requests.get(
API_URL,
headers=headers,
params={"offset": offset, "limit": PAGE_SIZE},
timeout=(10, 60),
)
response.raise_for_status()
payload = response.json()
if not isinstance(payload, dict) or not isinstance(payload.get("items"), list):
raise ValueError("Expected an object containing an items list")
items = payload["items"]
if not items:
break
results.extend(normalize(item) for item in items)
reported_total = payload.get("total")
returned_offset = payload.get("offset", offset)
returned_limit = payload.get("limit", len(items))
if not isinstance(returned_offset, int) or not isinstance(returned_limit, int) or returned_limit <= 0:
raise ValueError("Invalid offset or limit in API response")
offset = returned_offset + returned_limit
if isinstance(reported_total, int) and offset >= reported_total:
break
return results
if __name__ == "__main__":
records = fetch_all()
for record in records:
print(record)
Set API_URL and API_TOKEN in the server environment before running the script, and install the requests package if it is not already available. If the API uses a cursor or a next URL instead of offsets, replace the loop’s pagination logic: send the returned cursor or follow the documented continuation link until none remains. Do not send both cursor and offset parameters unless the API explicitly supports that combination.
The sample prints normalized records; in a production scraper, replace that output with a database write or a durable file export. For large datasets, persist each page as it arrives rather than retaining every record in memory. Add an idempotent upsert or a deduplication key if runs can overlap or be repeated.
Equivalent request patterns in cURL and Node.js
These snippets show the same basic authenticated GET request. They deliberately do not guess the target service’s response schema or pagination parameters. Add the endpoint’s documented inputs and use its actual continuation mechanism.
cURL
curl --fail-with-body --get "$API_URL"
--header "Authorization: Bearer $API_TOKEN"
--data-urlencode "offset=0"
--data-urlencode "limit=100"
--output response.json
Node.js
const apiUrl = process.env.API_URL;
const token = process.env.API_TOKEN;
if (!apiUrl || !token) {
throw new Error("Set API_URL and API_TOKEN in the server environment");
}
const url = new URL(apiUrl);
url.searchParams.set("offset", "0");
url.searchParams.set("limit", "100");
const response = await fetch(url, {
headers: { Authorization: `Bearer ${token}` },
signal: AbortSignal.timeout(60000),
});
if (!response.ok) {
throw new Error(`API request failed: HTTP ${response.status}`);
}
const payload = await response.json();
if (!payload || !Array.isArray(payload.items)) {
throw new Error("Expected an items array in the API response");
}
const records = payload.items.map((item) => {
if (item.id == null) throw new Error("Record is missing its id");
return { id: String(item.id), name: String(item.name ?? "") };
});
console.log(records);
For a POST endpoint, consult its contract for the correct content type and body shape. Do not convert a GET to POST, or vice versa, merely to make a parameter fit. Likewise, headers and cookies may affect what the target returns, but should be added only when documented and permitted.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Handle asynchronous jobs, pagination, and failures
Hosted jobs are a multi-step workflow
Some platforms do not return scraped data from the request that starts the work. A common pattern is to discover a tool, create a run, poll a status endpoint, and export the completed run’s dataset. Treat job creation and data retrieval as separate operations: preserve the run identifier, use the status response to determine when it is complete, and retrieve the dataset only through the documented export endpoint. Set a reasonable polling interval and overall deadline rather than polling continuously.
Pagination requires a definite stopping condition
Follow the API’s actual continuation contract. With offset-and-limit pagination, advance using the returned values when provided, and stop when the page is empty or the offset reaches the reported total. With cursor pagination, pass the returned cursor exactly as documented and stop when no next cursor or link remains. Verify that the code cannot repeat the same cursor forever. For changing datasets, offset pages can shift between requests; if the provider offers stable cursors or snapshot tokens, prefer those for consistent exports.
Throttle and retry with bounds
Check the target’s quota and concurrency rules before increasing parallel requests. api.data.gov’s documentation, accessed in 2026, states that participating services have a default limit of 1,000 requests per hour, with service-specific variation; excess requests can receive HTTP 429. That is a documented default for those services, not a universal API limit. Respect a service’s rate-limit headers and any Retry-After instruction.
Use bounded exponential backoff with jitter for transient network failures, HTTP 429 responses, and selected 5xx responses. Set a maximum number of attempts and an overall time budget. Do not retry ordinary 4xx errors such as invalid input or rejected credentials without first correcting the cause. Make writes idempotent where possible so a retry does not create duplicate records.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsRank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Latency depends on the endpoint and whether it must render a page or route through a proxy. ScraperAPI’s FAQ, accessed in 2026, describes typical latency of roughly 4–12 seconds for its requests and says some can take up to 60 seconds; that is the vendor’s operational guidance, not an independent benchmark or a promise for another provider. Set timeouts based on the endpoint’s documented behavior and your application’s deadline.
Choose between a data API and rendered-page scraping
| Approach | Use it when | Main work you still own |
|---|---|---|
| Direct data API | The data owner exposes the fields you need in a documented endpoint. | Authentication, pagination, schema mapping, quota handling, and persistence. |
| Rendered-page API | The required content appears only after browser-side JavaScript runs, or the target offers no usable data API. | Choosing rendering and request options, parsing returned HTML, and handling provider limits and changing page structure. |
| Hosted scraper platform | A maintained scraper or managed run workflow removes substantial browser, proxy, CAPTCHA, scheduling, or storage work. | Provider-specific run lifecycle, result schema, export, pricing, quota, and concurrency integration. |
A documented data API is often easier to parse because it returns structured fields, but it does not remove the need to handle authentication, pagination, or contract changes. Rendered-page services can help with JavaScript-created content; WebScraping.AI documents JavaScript execution and custom headers, while ScraperAPI documents JavaScript rendering and proxy options. A hosted platform can reduce operational work, but the provider’s supported targets, behavior, and output schema become dependencies. Compare request and parser control, synchronous versus asynchronous operation, pagination, rendering, geotargeting, retry behavior, export formats, and metering before migrating.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Validate the modified scraper before production
Save representative API responses as fixtures and test the parser without making live requests. Include a normal page, an empty page, a record with a missing optional field, malformed data, and each documented error format. Also test authentication failures, throttling, and a server error so that the program does not accidentally interpret an error object as a valid record.
- Confirm the request method, URL encoding, headers, and body against the provider’s current documentation.
- Check that each expected field has the intended type, including IDs, prices, dates, and booleans.
- Verify that pagination terminates, does not skip or repeat a page, and handles a total that changes during a run.
- Ensure one malformed record is reported clearly and does not silently corrupt the rest of the dataset.
- Test retries with a fixed attempt limit; verify that permanent client errors are surfaced rather than retried forever.
- Check logs, saved fixtures, and exported files for tokens, cookies, or other sensitive values.
Monitor record counts, failed requests, retry counts, and run duration after deployment. A sudden change can indicate an API schema change, a tightened quota, a changed target page, or a parser assumption that no longer holds. Keep the raw response and normalized output distinguishable so you can identify whether a problem occurred during retrieval or transformation.
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Troubleshooting common API-scraping problems
| Symptom | Likely cause | What to check or change |
|---|---|---|
| HTTP 401 or 403 | Missing, invalid, expired, or incorrectly placed credentials; the account may not be allowed to use the endpoint. | Check the documented auth scheme, token scope and expiry, account access, and whether the secret is actually present in the server environment. Do not put a secret in a public client. |
| HTTP 400 or a validation error | Wrong parameter name, encoding, method, or request-body shape. | Compare the request with the endpoint’s current schema. Inspect the returned error body without logging secrets. |
| HTTP 429 | Rate or concurrency limit reached. | Reduce request rate, observe rate-limit headers and any retry delay, and check whether the service has a separate quota for the credential or endpoint. |
| HTTP 5xx, timeout, or intermittent failure | Temporary provider or network issue; a rendered-page request may also exceed the chosen timeout. | Use a longer but bounded timeout where appropriate, retry transient failures with capped backoff, and track the request or run identifier if supplied. |
| Successful response but no records | Wrong endpoint or query, an empty result, an asynchronous job not yet complete, or a parser looking for the wrong property. | Inspect the raw status, body, and headers. Confirm whether the endpoint returns items, another records key, or a run ID that requires polling. |
| Duplicate or missing records | Pagination logic uses the wrong cursor or offset, page boundaries shift, or retries repeat a write. | Follow the response continuation values, log page identifiers, use a stable key for deduplication, and make storage writes idempotent. |
| Data differs from what the browser shows | The API endpoint may expose different data, require documented headers or session context, or omit JavaScript-rendered content. | Check the endpoint’s intended audience and response contract. If the needed content is page-rendered, use an appropriate rendering workflow rather than assuming JSON and browser output are interchangeable. |
Cost, reliability, and permission checks
Before deployment, estimate calls per run, pages per dataset, reruns, retries, and concurrent jobs. Compare those figures with the provider’s current quota and metering rules; a request-based service and a per-result platform may count work differently. Include the cost of rendering, proxy use, and storage when those are separately metered. Do not treat a vendor’s latency or success statement as an independent guarantee: for example, WebScraping.AI documents an 80%+ success rate for most websites, a vendor claim rather than a universal or independently verified result.
Reliability comes from bounded work and observable outcomes: timeouts, capped retries, explicit pagination stops, idempotent persistence, and alerts for unexpected empty results. Recheck current API terms, authentication rules, robots and website terms, and the intended use of the data before deployment. Access to an endpoint or a successful scrape does not itself establish permission to collect or reuse a particular site’s data.
Or skip the browser setup
If the task is to capture a page as an image or PDF—not extract structured records—ScreenshotNeo is a screenshot API and MCP server made by Yorker Media. Its clean-shot flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
One GET request returns an image or PDF. Example cURL call (see the ScreenshotNeo API documentation for options):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots, and every feature is on every plan. For AI agents, the MCP server avoids writing a separate browser-capture integration. This is a screenshot workflow, not a substitute for a data API when your output needs structured records. Sign up for the free plan: 1,000 screenshots a month, no card required.
Frequently Asked Questions
Does using an API mean I no longer need to parse data?
No. An API changes how the data is requested and may provide a structured response, but your application still needs to interpret, validate, and store the fields it receives.
Can I use a screenshot API to extract structured product records?
A screenshot API returns a visual capture, not a structured product dataset. Use a data endpoint or a scraping workflow that returns parseable HTML or JSON for record extraction.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →

