What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
There is no universal “scrape every page” command. Build an asynchronous workflow that (1) discovers the next page or scroll state, (2) waits for the target site’s content to be ready, (3) extracts records with resilient locators, (4) records what was processed, and (5) stops on a site-specific end condition. The templates below show how to do that for numbered pagination and infinite scrolling without assuming that a browser’s load event means the application is finished.
What “all pages” means in an async scraper
In Playwright discussions, “pages” can mean browser tabs managed by one context or paginated result states. This article uses result pages: every numbered or “Next” state, and every batch appended by infinite scrolling. A Playwright Page object is the browser tab that displays one of those states.
Before writing code, define four site-specific facts:
- The starting URL and the fields you need.
- How another result state is reached: a link, button, URL parameter, or scroll gesture.
- What proves that the current records are ready: a visible result, a loading indicator disappearing, a count reaching an expected value, or an end marker.
- What proves that collection is complete: a disabled/absent next control, an end-of-results marker, or no measurable increase after a bounded attempt.
These conditions are part of the target site’s behavior. A selector or sleep interval that works on one site is not a general rule.
#1 Best Overall
Install Playwright and prepare an async project
Install the Python package and browser binaries in your environment:
python -m pip install playwright
python -m playwright install chromium
The async API requires an event loop. The following imports and helper show a practical baseline with a context, one page, and explicit timeouts. Keep the browser context alive while you process related result states so cookies and other session state are preserved.
import asyncio
import json
from pathlib import Path
from playwright.async_api import async_playwright, TimeoutError as PlaywrightTimeoutError
START_URL = "https://example.com/products"
OUT = Path("products.jsonl")
async def main():
async with async_playwright() as pw:
browser = await pw.chromium.launch(headless=True)
context = await browser.new_context()
page = await context.new_page()
page.set_default_timeout(15_000)
try:
records = await process_listing(page, START_URL)
with OUT.open("w", encoding="utf-8") as f:
for record in records:
f.write(json.dumps(record, ensure_ascii=False) + "n")
finally:
await context.close()
await browser.close()
if __name__ == "__main__":
asyncio.run(main())
Use a real URL and selectors from the site you are allowed to access. Do not treat browser automation as permission to bypass access controls.
Wait for the application, not just the network
A navigation can resolve while JavaScript is still rendering records. The load event only describes a browser lifecycle milestone; it does not guarantee that the data your extractor needs exists. Wait for a meaningful condition such as the first result becoming visible or a site-specific loading element disappearing.
async def wait_for_results(page):
# Replace these selectors with stable ones from the target site.
await page.locator('[data-testid="result-card"]').first.wait_for(state="visible")
await page.locator('[data-testid="results-spinner"]').wait_for(state="hidden")
If the site exposes a result count, wait for that count instead. Avoid a fixed delay as your only readiness check: it can be too short on a slow run and wasteful on a fast one.
Rank #2
Extract records with resilient locators
Locators are evaluated when used and provide Playwright’s auto-waiting and retry behavior. Prefer user-facing roles, labels, text, or explicit test IDs over selectors tied to incidental nesting or generated class names. A stable card extractor might look like this:
async def extract_current_records(page):
cards = page.locator('[data-testid="result-card"]')
records = []
count = await cards.count()
for index in range(count):
card = cards.nth(index)
title = (await card.get_by_role("heading").inner_text()).strip()
link = await card.get_by_role("link").first.get_attribute("href")
price = await card.locator('[data-testid="price"]').inner_text()
records.append({"title": title, "url": link, "price": price.strip()})
return records
locator.all() is useful only after the set is stable. It returns the matches present immediately; it does not wait for a dynamic list to finish changing. Calling it while cards are still arriving can produce incomplete or flaky results. Waiting first and then using count()/nth(), as above, makes the timing explicit.
Process numbered or “Next” pagination
For conventional pagination, process the current URL, record a visited key, then activate the site’s actual next mechanism. A link-based implementation follows:
async def process_listing(page, start_url):
records = []
seen_states = set()
url = start_url
while url not in seen_states:
seen_states.add(url)
await page.goto(url, wait_until="domcontentloaded")
await wait_for_results(page)
records.extend(await extract_current_records(page))
next_link = page.get_by_role("link", name="Next")
if await next_link.count() == 0:
break
if await next_link.is_disabled():
break
href = await next_link.get_attribute("href")
if not href:
break
url = await page.evaluate("(href) => new URL(href, location.href).href", href)
return records
Some sites render a button rather than a link. In that case, click it and wait for a measurable state change—such as the URL changing, a page-number label changing, or the first card’s text changing—before extracting again:
async def advance_button_pagination(page):
current_marker = await page.locator('[aria-current="page"]').inner_text()
next_button = page.get_by_role("button", name="Next")
if await next_button.count() == 0 or not await next_button.is_enabled():
return False
await next_button.click()
await page.wait_for_function(
"marker => document.querySelector('[aria-current="page"]')?.textContent.trim() !== marker",
current_marker,
)
await wait_for_results(page)
return True
If the site reuses the same URL for every state, track a page number or a unique record identifier rather than relying on page.url. Always keep a visited set; it protects you from malformed “next” links that loop back to an earlier state.
Handle infinite scrolling safely
Infinite lists require a bounded loop. Scroll the relevant container (or the document), wait for the number of records to increase, and stop when the site exposes an end marker or repeated attempts produce no new records.
async def process_infinite_list(page, start_url, max_rounds=200):
await page.goto(start_url, wait_until="domcontentloaded")
await wait_for_results(page)
cards = page.locator('[data-testid="result-card"]')
seen_urls = set()
records = []
stagnant_rounds = 0
for _ in range(max_rounds):
before = await cards.count()
for record in await extract_current_records(page):
key = record.get("url") or record.get("title")
if key not in seen_urls:
seen_urls.add(key)
records.append(record)
end_marker = page.locator('[data-testid="end-of-results"]')
if await end_marker.count() and await end_marker.is_visible():
break
await cards.last.scroll_into_view_if_needed()
try:
await page.wait_for_function(
"before => document.querySelectorAll('[data-testid="result-card"]').length > before",
before,
timeout=10_000,
)
stagnant_rounds = 0
except PlaywrightTimeoutError:
stagnant_rounds += 1
if stagnant_rounds >= 2:
break
return records
If the list is inside a scrollable element, scroll that element instead of the window:
Recommended Free Tools
container = page.locator('[data-testid="results-panel"]')
await container.evaluate("el => el.scrollTop = el.scrollHeight")
Use the site’s own loading signal when possible. A spinner disappearing, a “showing 40 of 200” label changing, or an end marker is more reliable than waiting an arbitrary number of seconds. Keep both a maximum-round limit and a no-progress limit so a broken endpoint cannot run forever.
Scrape detail pages discovered from each result
After collecting list records, visit detail URLs in a controlled sequence or with a small worker pool. A browser context can host multiple pages, but opening every URL at once consumes memory and increases load on the target. There is no universal safe concurrency number; choose a conservative bound for the site and machine, then adjust based on observed failures.
async def scrape_detail(context, record):
detail = await context.new_page()
try:
await detail.goto(record["url"], wait_until="domcontentloaded")
await detail.locator("main").wait_for(state="visible")
record["description"] = (await detail.locator("main").inner_text()).strip()
return record
finally:
await detail.close()
async def scrape_details_bounded(context, records, workers=4):
semaphore = asyncio.Semaphore(workers)
async def run(record):
async with semaphore:
try:
return await scrape_detail(context, record)
except Exception as exc:
return {**record, "error": str(exc)}
return await asyncio.gather(*(run(r) for r in records))
Persist each successful record as you go (JSON Lines, a database, or another durable store). Keep failed URLs and exception messages in a separate field or file so one timeout does not silently invalidate the entire collection. If you resume later, load the completed keys and skip them.
Pagination versus infinite scroll: choose the matching loop
| Decision | Numbered/Next pagination | Infinite scrolling | Implementation consequence |
|---|---|---|---|
| State transition | URL, page number, or next control | Scroll gesture that appends records | Use visited URLs/page IDs versus record-count growth. |
| Completion | Next absent or disabled | End marker or repeated no-progress attempts | Both require a defensive upper bound. |
| Readiness | New page’s results visible | New batch rendered after scrolling | Wait on content or application state, not only navigation/load. |
Troubleshooting common failures
Only the first batch is returned
The list was read before the application finished appending records, or locator.all() was called while it was changing. Wait for a count increase or a loading indicator to disappear, then extract.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match“Next” clicks but records repeat
The click did not change the state, or the site requires a different control. Capture a page-number/URL marker before clicking and wait until that marker changes. Maintain a visited set to stop loops.
A timeout occurs although the page looks loaded
The chosen readiness selector may not exist for an empty result, an error state, or a different layout. Add explicit branches for “no results” and site error messages, and inspect the DOM for a stable role, label, or test ID.
Infinite scrolling stops prematurely
The script may be scrolling the wrong element, waiting on a count that includes hidden templates, or hitting a rate limit. Scroll the actual results container, measure visible record identifiers, and use the site’s loading/end signal.
Duplicate detail pages appear
Normalize and deduplicate canonical URLs (or another stable record ID) before scheduling detail work. Keep the original URL for auditability.
Best Value
The run becomes unreliable at higher concurrency
Reduce the worker bound, reuse one context, add per-page timeouts, and record failures for retry. Concurrency improves throughput only when the site and your machine can sustain it; official Playwright documentation establishes that contexts can contain multiple pages, not a universal concurrency limit.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Performance, reliability, and responsible operation
- Bound work: set maximum pagination rounds, scroll rounds, and per-navigation timeouts.
- Measure progress: log URL/page ID, records found, elapsed time, and failure reason after each state.
- Retry narrowly: retry transient navigation or network failures with a small limit; do not endlessly retry deterministic selector errors.
- Preserve state: use one context for a session that needs cookies, but close detail pages promptly.
- Respect the site: follow terms, robots guidance where applicable, authentication rules, privacy obligations, and rate limits. Automation does not grant permission to collect data.
- Validate output: check required fields, unique keys, and expected page counts before publishing or loading results downstream.
No general throughput or speedup figure applies to every site. Browser rendering cost, JavaScript behavior, network latency, throttling, and the selected concurrency all change the result, so benchmark your own defined workload if performance matters.
Or skip the browser setup
If you only need clean screenshots or PDFs of pages rather than DOM-level extraction, ScreenshotNeo provides a single GET request. It accepts cookie/consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Full request options and parameter names are documented at https://screenshotneo.com/docs/.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every feature is included on every plan. The Free plan provides 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.
FAQ
Should I use one browser page for every result page?
Reuse one page for sequential pagination. For independent detail URLs, use a small bounded pool of pages and close each page after extraction.
Is networkidle always the best wait condition?
No. Analytics, ads, and long-lived connections can prevent network-idle from being reached. A selector or application-state condition tied to the records you need is usually more meaningful.
How do I know whether to scroll the window or a panel?
Inspect which element actually has a scrollable height and receives the site’s wheel/keyboard events. Scroll that element and observe a measurable increase in record identifiers.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

