Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Use Playwright actions in the Scrapy request metadata, then wait for a DOM change that proves the scroll loaded new content. For one scroll, add PageMethod objects to playwright_page_methods. For an infinite list, repeat the scroll-and-wait cycle until the item count stops increasing or the page exposes a terminal signal. If the scrollbar belongs to a feed or modal rather than the document, scroll that element’s scrollTop or send wheel input to it.
Prerequisites and architecture
scrapy-playwright is a Scrapy download handler that performs requests with Playwright while retaining Scrapy’s scheduling and item-processing workflow. The project’s documented minimum versions at the time of writing are Python 3.10, Scrapy 2.7 and Playwright 1.40; verify the current repository requirements before pinning them because they can change.
Install the integration and the browser binaries:
python -m pip install scrapy-playwright
playwright install
Enable the handler in your Scrapy settings (the exact download-handler path is shown in the current scrapy-playwright documentation), then opt individual requests into browser rendering with meta={"playwright": True}. This keeps browser work inside Scrapy instead of bypassing middleware, scheduling and duplicate filtering by driving Playwright as a separate crawler.
Scroll the document once and prove that content arrived
The smallest useful pattern is a request with three page methods: wait for the initial item, scroll the document, and wait for a selector representing the next item. The second selector is the proof of progress; a fixed sleep only proves that time passed.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
import scrapy
from scrapy_playwright.page import PageMethod
class QuotesSpider(scrapy.Spider):
name = "quotes"
def start_requests(self):
yield scrapy.Request(
url="https://quotes.toscrape.com/scroll",
meta={
"playwright": True,
"playwright_page_methods": [
PageMethod("wait_for_selector", "div.quote"),
PageMethod(
"evaluate",
"window.scrollBy(0, document.body.scrollHeight)",
),
PageMethod(
"wait_for_selector",
"div.quote:nth-child(11)",
),
],
},
callback=self.parse,
)
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
}
PageMethod actions run before the browser response is handed to parse. The first wait avoids racing the initial render. window.scrollBy sends the page downward. The final wait targets an item that was not present in the initial batch, so the callback sees evidence that the list advanced.
A selector for the first card, such as div.quote, is not a progress check: that card already existed before scrolling. Replace the example’s item number with a selector appropriate for the site you crawl.
Repeat scrolling for an infinite list
For several rounds, use a callable page method. It can inspect the page, perform one scroll, and wait for the next sentinel. There is no universal iteration count or delay: stop according to the target site’s behavior and your extraction goal.
from scrapy_playwright.page import PageMethod
async def scroll_page(page):
previous_count = await page.locator("article.card").count()
await page.evaluate("window.scrollBy(0, document.body.scrollHeight)")
try:
await page.wait_for_function(
"([selector, oldCount]) => "
"document.querySelectorAll(selector).length > oldCount",
["article.card", previous_count],
timeout=10_000,
)
except Exception:
# The caller can treat an unchanged count as the end of the feed.
return False
return True
Use it from a request:
yield scrapy.Request(
"https://example.com/feed",
meta={
"playwright": True,
"playwright_page_methods": [
PageMethod("wait_for_selector", "article.card"),
PageMethod(scroll_page),
],
},
)
For a complete crawler, repeat requests or page actions until one of these bounded conditions is true:
Recommended Free Tools
- The number of cards is unchanged after a scroll-and-wait attempt.
- A “no more results” or equivalent terminal element becomes visible.
- The next-page control is missing or disabled.
- You have collected the known number of records required by the job.
- A safety limit you chose for this crawl is reached.
When a site appends cards in batches, wait for a count increase or a new card with a stable identity. Do not assume that a scroll event itself means an API response completed.
Scroll a nested feed, panel or modal
Many interfaces keep the document fixed while an inner element owns the scrollbar. Scrolling window in that situation changes nothing useful. First identify the element whose computed layout has an overflowing height, such as div.feed or [role="dialog"] .results.
Use wheel input
async def wheel_inner_feed(page):
feed = page.locator("div.feed")
await feed.wait_for()
await feed.hover()
await page.mouse.wheel(0, 900)
await page.wait_for_selector("div.feed article.card:nth-child(21)")
Hovering places the pointer over the intended scrolling region. The amount is pixels, not an item count; choose a value that matches the interface and then wait for a new card or another progress signal.
Change the element’s scrollTop directly
async def scroll_inner_element(page):
feed = page.locator("div.feed")
before = await feed.locator("article.card").count()
await feed.evaluate("(el) => { el.scrollTop += el.clientHeight; }")
await page.wait_for_function(
"([el, oldCount]) => el.querySelectorAll('article.card').length > oldCount",
[await feed.element_handle(), before],
)
A locator-level evaluation is useful when wheel events are intercepted or when you need deterministic element scrolling. In production code, prefer a locator evaluation that does not retain a stale element handle longer than necessary; after each scroll, re-read the locator and its count.
Rank #3
Find the real scroll owner
- Inspect the element containing the cards, not only
body. - Look for a fixed height and
overflow: autooroverflow: scroll. - Confirm that its
scrollHeightexceeds itsclientHeight. - Scroll that element, then wait for a card, count or terminal marker to change.
Choose locators that survive redesigns
Playwright recommends user-facing locators first: accessible roles, visible text and labels. For example:
page.get_by_role("article")
page.get_by_role("button", name="Load more")
page.get_by_text("No more results")
Use CSS or XPath when the page offers no stable accessible contract, or when a structural relationship is the only reliable signal. Keep selectors specific enough to identify a newly loaded item rather than an element that was already present. If cards have IDs or data attributes, a stable data-* selector is generally preferable to a deeply nested XPath.
Expose and close the Playwright page deliberately
Most jobs need no page handle: scrapy-playwright executes the declared page methods and returns a normal Scrapy response. Set playwright_include_page=True only when callback code must inspect the page, take a screenshot, perform additional actions or run custom JavaScript.
async def parse(self, response):
page = response.meta["playwright_page"]
try:
await page.get_by_role("button", name="Load more").click()
await page.get_by_role("article").last.wait_for()
yield {"html": await page.content()}
finally:
await page.close()
Always close an included page, including error paths. Leaving pages open consumes browser resources and can eventually stall the crawl. If you follow links or schedule additional requests from a callback, close the current page before returning unless the integration’s documented lifecycle explicitly transfers ownership.
Waiting strategies: what proves progress?
Wait for a new selector
This is the direct pattern used in the project README: after scrolling, wait for a selector such as the next card in the sequence. It is simple and efficient when the markup has predictable item order.
Wait for a count increase
Count the cards before the scroll and wait until the count is greater. This handles variable batch sizes and avoids hard-coding an item number.
Wait for a terminal state
Some feeds do not append cards when exhausted. Wait for a “no more results” marker, a disabled button or a network-idle state combined with an unchanged count. Network idle alone is not proof that data arrived: a page can become idle while returning an empty batch.
Use timeouts as failure handling, not success criteria
A timeout tells you that the expected signal did not appear. Treat it as a branch: retry if the site is transient, record the final count if the list is exhausted, or fail the item when missing content would make the result incomplete.
Best Value
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| No new cards after scrolling | The feed, not window, owns the scrollbar. |
Scroll the nested locator with mouse.wheel or scrollTop. |
| The wait succeeds immediately | You waited for an element that already existed. | Wait for a new card, an increased count or a terminal marker. |
| Only the first batch is scraped | The crawl ends after one scroll. | Loop with a bounded stopping condition. |
| Intermittent timeout | Loading time varies or the selector is unstable. | Use a meaningful selector, an appropriate timeout and a retry policy; avoid replacing the signal with a blind sleep. |
| Browser resources grow without bound | Included pages are never closed. | Close response.meta["playwright_page"] in a finally block. |
| Actions fail before the callback | Browser or package versions do not meet current requirements. | Check the scrapy-playwright project’s current documented minimums and rerun Playwright’s browser installation. |
| Cards appear visually but fields are empty | Content is rendered after the selector appears, or data is inside a shadow component. | Wait for the field or a stable card state, then inspect the rendered DOM with the page handle. |
Performance, reliability and crawl safety
- Use the smallest viewport and browser context settings that still trigger the site’s loading behavior.
- Wait on progress signals instead of adding long fixed delays; this reduces idle time on fast responses while preserving correctness on slow ones.
- Keep each scroll bounded. Infinite feeds can otherwise run forever, consume memory and repeatedly request content you do not need.
- Deduplicate items by a stable URL or ID because virtualized lists may recycle DOM nodes.
- Record the final item count and the reason the loop stopped. This makes an exhausted feed distinguishable from a failed wait.
- Respect the target site’s terms, robots policy and rate limits. Browser rendering does not remove those obligations.
Or skip the browser setup
If your goal is a clean image or PDF rather than a Scrapy extraction, ScreenshotNeo provides a single website-screenshot API request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not charged, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also provides an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
Use the API documentation at screenshotneo.com/docs/ for all options. A basic call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same request in Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element captures, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper sizes/margins/orientation/page ranges, HTML/CSS-to-image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for selectors/delays/network idle, blocking ads/trackers/requests/resource types, custom headers/cookies/user agents/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed public-image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work, easing migration.
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; Growth is $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to start.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePractical checklist
- Confirm the current Python, Scrapy, Playwright and browser requirements.
- Install
scrapy-playwrightand the Playwright browsers. - Set
meta["playwright"] = Trueon the request. - Add a
PageMethodthat scrolls the correct owner: the document or an inner locator. - Wait for a new selector, increased item count or terminal state.
- Repeat only while progress is demonstrated and your safety bound permits it.
- Use robust, user-facing locators before CSS or XPath.
- Include and close the page only when callback logic needs a Playwright handle.
Frequently Asked Questions
Does Playwright automatically scroll before every action?
Playwright usually scrolls an element into view before acting on it. That automatic behavior is different from deliberately scrolling an infinite list to trigger more data, which still requires an explicit scroll and a wait for progress.
Can I use a fixed sleep after scrolling?
You can, but it is weaker than waiting for a selector, count increase or terminal state. A sleep may be too short on a slow response and unnecessarily long on a fast one.
How do I know whether a page is exhausted?
Use a site-specific signal such as an unchanged item count, a visible end-of-results marker or a disabled/missing next-page control, together with a safety limit.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems

