Yes—you can scrape a JavaScript website with Python by driving a real browser with Selenium. This build-along tutorial creates a small, restartable scraper that waits for rendered content, uses maintainable locators, follows pagination, and writes structured JSON. It targets Selenium 4 with Python 3.10 or newer; current Selenium supports Chrome, Edge, Firefox, Safari, WebKitGTK and WPEWebKit.
What you will build
The example collects article cards from a fictional catalog page. Replace the URL and selectors with those from the site you are authorized to access. The scraper will:
- open a browser and navigate with
driver.get(); - wait for JavaScript-rendered cards instead of guessing with long sleeps;
- extract text, links and attributes;
- follow a “next” link until pagination ends;
- checkpoint results so a later run can resume safely; and
- always close the browser in a
finallyblock.
Selenium executes the page’s JavaScript, so it can see content that a plain HTTP request would receive only as an empty shell. That power costs more CPU, memory and time than an HTTP client, so use Selenium when browser execution is actually required.
1. Create the environment
- Install Python 3.10 or newer and verify it with
python --version. - Create and activate a virtual environment:
python -m venv .venv
macOS/Linux:source .venv/bin/activate
Windows PowerShell:.venvScriptsActivate.ps1 - Install Selenium:
python -m pip install -U selenium.
Modern Selenium uses Selenium Manager to locate and manage a compatible browser driver when you start webdriver.Chrome(). You normally do not need to download ChromeDriver yourself. Manual driver paths are still useful in locked-down machines, custom browser installations or when diagnosing a browser/driver mismatch.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute#1 Best Overall
2. Launch a browser deliberately
Start with a visible browser while developing. Headless mode is convenient in CI, but a visible window makes consent dialogs, redirects and selector mistakes easier to inspect.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
options = Options()
# Uncomment for CI or a server without a display:
# options.add_argument("--headless=new")
options.page_load_strategy = "normal"
driver = webdriver.Chrome(options=options)
driver.set_page_load_timeout(30)
driver.set_script_timeout(30)
try:
driver.get("https://example.com")
print(driver.title)
finally:
driver.quit()
Choose a page-load strategy
- normal waits for the load event and is the safest default.
- eager returns after DOMContentLoaded; it can reduce waiting when images and other assets are irrelevant.
- none returns without waiting for page loading. You must then synchronize every required state explicitly.
Page-load, script and implicit-wait timeouts are separate controls. Set each intentionally rather than relying on an accidental default. A proxy can be configured through browser options when a restricted network, traffic capture setup or mock backend requires one.
3. Inspect the rendered DOM
Load the page in a normal browser first. Use developer tools’ Elements panel after the content appears, not only the initial “view source.” Identify one card, its title link, and the pagination control. Check whether the site uses an iframe; elements inside an iframe require driver.switch_to.frame(...) before lookup and driver.switch_to.default_content() afterward.
Pick locators that survive redesigns
- Prefer a unique, stable HTML
id. Selenium’s guidance says: “In general, if HTML IDs are available, unique, and consistently predictable, they are the preferred method for locating an element on a page.” - If no reliable ID exists, use a compact CSS selector such as
article.card a.title. - Use XPath when you need relationships or text-based logic, but keep it narrow and readable. XPath is usually harder to debug and slower than a good CSS selector.
- Avoid generated IDs, deeply nested paths and classes that only describe visual styling.
4. Wait for the state you need
driver.get() returning proves only that the selected page-load milestone occurred. A single-page app may still be fetching data. Use an explicit wait for the element or state that means your data is ready.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Rank #2
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
wait = WebDriverWait(driver, 10)
container = wait.until(
EC.presence_of_element_located((By.CSS_SELECTOR, "article.card"))
)
wait.until(EC.visibility_of(container))
Presence means the node exists in the DOM; visibility additionally requires it to be displayed. For a specific application state, wait for a URL change, a button to become clickable, a loading spinner to disappear, or expected text to appear. A short fixed time.sleep() can be useful for a known animation, but it is not a substitute for a condition: it can under-wait on a slow run and waste time on a fast one.
Selenium also offers implicit waits, which apply to every element lookup. Do not mix implicit and explicit waits. One synchronization policy is easier to reason about, and mixing them can make explicit waits take much longer than their stated timeout.
5. Build the scraper
Save this as scrape_catalog.py. The selectors and URL are examples; replace them after inspecting your target site.
import json
import time
from pathlib import Path
from selenium import webdriver
from selenium.common.exceptions import (
StaleElementReferenceException,
TimeoutException,
WebDriverException,
)
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
START_URL = "https://example.com/catalog"
OUTPUT = Path("items.json")
CARD = (By.CSS_SELECTOR, "article.card")
NEXT = (By.CSS_SELECTOR, "a.next")
def text_or_empty(element, selector):
try:
return element.find_element(By.CSS_SELECTOR, selector).text.strip()
except Exception:
return ""
def collect_page(driver, wait):
cards = wait.until(EC.presence_of_all_elements_located(CARD))
rows = []
for card in cards:
try:
link = card.find_element(By.CSS_SELECTOR, "a.title")
rows.append({
"title": link.text.strip(),
"url": link.get_attribute("href"),
"summary": text_or_empty(card, ".summary"),
})
except StaleElementReferenceException:
# The framework re-rendered the card; skip it and let the next run retry.
continue
return rows
def save_checkpoint(rows):
OUTPUT.write_text(json.dumps(rows, ensure_ascii=False, indent=2), encoding="utf-8")
def main():
options = Options()
# options.add_argument("--headless=new")
options.page_load_strategy = "normal"
driver = webdriver.Chrome(options=options)
driver.set_page_load_timeout(30)
driver.set_script_timeout(30)
wait = WebDriverWait(driver, 15)
rows = []
seen_pages = set()
try:
driver.get(START_URL)
for page_number in range(1, 101):
page_key = driver.current_url
if page_key in seen_pages:
break
seen_pages.add(page_key)
rows.extend(collect_page(driver, wait))
save_checkpoint(rows)
try:
next_link = wait.until(EC.presence_of_element_located(NEXT))
except TimeoutException:
break
if not next_link.is_enabled():
break
old_url = driver.current_url
driver.execute_script("arguments[0].click();", next_link)
wait.until(lambda d: d.current_url != old_url)
wait.until(EC.presence_of_all_elements_located(CARD))
time.sleep(0.2) # only to let a brief card animation settle
except (TimeoutException, WebDriverException) as exc:
print(f"Stopped after saving a checkpoint: {exc}")
finally:
driver.quit()
if __name__ == "__main__":
main()
Run it with python scrape_catalog.py. The JSON file is written after every page, so a network failure does not discard earlier work. For a production crawler, add a resume key (for example, the last URL), deduplicate by canonical URL, and persist a retry queue rather than restarting blindly.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →6. Extract more than visible text
Attributes and links
Use get_attribute("href"), get_attribute("data-id") or another attribute instead of parsing HTML strings. For an image URL, inspect src and lazy-loading attributes such as data-src. If text is rendered only after scrolling, scroll the element or viewport and wait for the card count to increase.
Tables
Locate table tbody tr, then extract each cell in a fixed order. Validate the number of cells before indexing; responsive tables sometimes hide columns at narrow viewport widths.
Clicks, filters and infinite scroll
Wait for a clickable control, click it, then wait for a changed URL, changed result count or disappearance of a loading indicator. For infinite scroll, record the previous item count, scroll, and wait until the count increases. Stop after several scrolls with no increase to avoid an endless loop.
7. Sessions, pagination and reliability
Keep one driver alive when a site requires cookies or a login session. Use a dedicated test account, never hard-code credentials, and store cookies only where your security policy permits. Detect the end of pagination by a missing or disabled next control, an unchanged URL, or a repeated page signature. Cap page counts and retries.
Free tools Windows power users keep installed
One-click scans. No signup required.
- Transient failures: retry a navigation a small, capped number of times with increasing delays.
- Stale elements: locate the element again after a framework re-render instead of reusing the old reference.
- Slow pages: increase the explicit wait only for the known slow condition; do not set every timeout to several minutes.
- Large jobs: checkpoint frequently, log URL and page number, and write records incrementally.
- Parallelism: multiple browsers multiply memory and request load. Start serially, then add bounded workers only when the site and machine can handle it.
8. Responsible scraping
Read the site’s terms and access rules before automating. Inspect robots.txt; RFC 9309 defines the Robots Exclusion Protocol. Robots rules are an access signal, not a blanket legal determination, so obtain permission where required and stop when a site blocks automation. Identify your user agent when appropriate, respect rate limits, avoid bypassing bot checks or access controls, and collect only the personal data you genuinely need. A technically successful request can still violate a contract, privacy rule or local law.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server when you need a rendered page image or PDF rather than a custom data-extraction workflow. One GET request returns PNG, JPEG, WebP or PDF. It accepts cookie/consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients.
cURL (see the ScreenshotNeo documentation):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Every plan includes its capture options: full-page lazy-image loading, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper/margins/landscape/page ranges, HTML/CSS rendering, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, request and resource blocking, headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage API and OpenAPI specification. Familiar parameter names from other screenshot APIs are accepted to ease migration.
| Plan | Included shots | Price |
|---|---|---|
| Free | 1,000/month | $0, no card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
Troubleshooting Selenium failures
“Element not found”
The selector may be wrong, the element may be inside an iframe, or JavaScript may not have rendered it yet. Verify the live DOM, switch into the correct frame, and wait for the target condition.
Best Value
“Element is not clickable”
A cookie dialog, overlay or off-screen element may intercept the click. Wait for clickability, scroll it into view, close the overlay if permitted, and confirm that the element is enabled.
Driver or browser version mismatch
Update Selenium and the browser, let Selenium Manager resolve the driver, and check that a corporate policy is not forcing an incompatible executable. Use an explicit driver path only when you control the matching versions.
TimeoutException
Capture the current URL and a screenshot, then determine whether the site is slow, blocked, redirected to login, or returning a different layout. Increase only the relevant wait and add a bounded retry.
Recommended Free Tools
Headless differs from headed mode
Set a realistic window size, wait for visibility rather than coordinates, and compare user-agent, viewport and locale settings. Some sites deliver different markup to headless browsers; do not assume the two modes are identical.
Frequently Asked Questions
Can Selenium scrape content behind a login?
It can navigate an authorized session, but use a permitted test account, protect credentials and follow the site’s terms and applicable law.
When should I use requests instead of Selenium?
Use an HTTP client when the data is present in the response or a documented API; choose Selenium when JavaScript execution, clicks, scrolling or browser cookies are necessary.
Is robots.txt permission to scrape?
No. It communicates access preferences; it does not by itself settle contractual or legal permission.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

