Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Pyppeteer to launch Chromium, open the page, and read the first matching div with querySelectorEval. For several matching divs, use querySelectorAllEval. The examples below include Linux setup, waiting for dynamically rendered content, handling missing selectors, and closing Chromium reliably. One maintenance caveat: Pyppeteer’s own project documentation describes it as unmaintained and suggests evaluating Playwright Python for new projects.

Extract text from one div

Install Pyppeteer, launch a headless browser, navigate to the target URL, and evaluate textContent on the element selected by a CSS selector. This complete asynchronous example returns the trimmed text of the first matching element:

import asyncio
from pyppeteer import launch

async def extract_div_text(url: str, selector: str) -> str:
    browser = await launch(headless=True)
    try:
        page = await browser.newPage()
        await page.goto(url, {"waitUntil": "networkidle2"})
        return await page.querySelectorEval(
            selector,
            "node => node.textContent.trim()"
        )
    finally:
        await browser.close()


if __name__ == "__main__":
    text = asyncio.get_event_loop().run_until_complete(
        extract_div_text("https://example.com", "div.article")
    )
    print(text)

Replace https://example.com with the page you are authorized to access, and div.article with a selector that identifies the div you need. querySelectorEval evaluates the supplied function against the first matching element. If nothing matches, it raises an error rather than returning an empty string. The finally block closes Chromium whether navigation and extraction succeed or raise an exception.

Install Pyppeteer and prepare Chromium on Linux

  1. Install the Python package: run python3 -m pip install pyppeteer. The Pyppeteer repository README and project guide document this installation command.
  2. Download Chromium in advance if needed: Pyppeteer may download Chromium on first use. To trigger that download before running your script, run pyppeteer-install.
  3. Run the script as the same user that will run your application: the documented default Linux download directory is /home/<username>/.local/share/pyppeteer. If XDG_DATA_HOME is set, the documented location is $XDG_DATA_HOME/pyppeteer.

Pyppeteer documentation and its repository README give different approximate first-run download estimates—about 100 MB and about 150 MB, respectively. Those are version-era documentation estimates, not a guaranteed current download size. Allow for a Chromium download and check which data directory applies to the user and environment running your program.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The project describes Pyppeteer as an unofficial Python port of Puppeteer and says its repository is unmaintained. That makes it important to pin and validate the Python and Chromium versions for a production workflow. For a new scraper, assess whether a maintained browser-automation library such as Playwright Python is a better fit; Pyppeteer remains an option when compatibility with an existing script or environment is the priority.

Choose the right selector and extraction method

One div: querySelectorEval

Use this when the CSS selector should identify a single element and you want a concise result. It selects only the first match, so a selector such as div.article does not automatically return every article div on the page.

One div: an element handle plus evaluate

If you want to separate finding the element from reading it, retrieve an element handle first and pass it to page.evaluate:

element = await page.querySelector("div.article")
if element is None:
    raise LookupError("No element matched div.article")

text = await page.evaluate("(element) => element.textContent", element)
print(text.strip())

This makes the no-match check explicit before evaluation. The Pyppeteer project guide documents the element-plus-evaluate pattern for reading text. Trim the returned value if leading and trailing whitespace is not useful to your application.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Every matching div: querySelectorAllEval

Use querySelectorAllEval when the result should include all matches. The callback receives the matched nodes and maps each one to its trimmed text:

texts = await page.querySelectorAllEval(
    "div.article",
    "nodes => nodes.map(node => node.textContent.trim())"
)

for text in texts:
    print(text)

The documented all-matches method returns the results of evaluating over matching elements; with this callback, texts is a list of strings. Unlike the single-element evaluation, a selector with no matches can be handled as an empty result, so decide whether that is acceptable or should count as an application error.

Use a selector that identifies the intended content

Prefer a stable ID, class, or data attribute that belongs to the content you want. A broad selector such as div is likely to match a page container or many unrelated elements; a more specific selector such as div.article narrows the target. The selector is evaluated in the page’s document, so check it against the actual page markup rather than assuming a class name from a different layout or URL.

Wait for content that appears after navigation

The example waits for networkidle2 during navigation, then immediately queries the selector. That works when the content is present by the time navigation reaches that condition. Some sites populate a div later through client-side code, so the page can be navigable before the target element exists. In that case, wait for the specific selector before extracting it:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
await page.goto(url, {"waitUntil": "networkidle2"})
await page.waitForSelector(selector)
text = await page.querySelectorEval(
    selector,
    "node => node.textContent.trim()"
)

Use the actual target selector in waitForSelector; waiting for a general page condition does not establish that a particular content block has rendered. The appropriate wait depends on the page. If the selector never appears, inspect the page URL and markup, confirm that the target is not inside a different frame, and decide how your script should handle a timeout or absent element.

textContent reads text from the selected DOM element. It is a straightforward choice when the task is to extract the element’s text content. If your requirement is specifically the text as presented visually on the page, verify the result against that page: layout, whitespace, hidden content, and site-specific rendering can affect whether a raw DOM text value meets that requirement. Pyppeteer examples demonstrate both textContent and innerText, but the extraction choice should follow the result your downstream task actually needs.

Handle missing selectors and common failures

Symptom Likely cause What to check or change
querySelectorEval raises because no element matches The selector does not match the current document, or the target has not rendered yet. Check the selector against the page, wait for the target selector when rendering is asynchronous, or use querySelector and explicitly handle a missing element.
The result is empty or lacks expected content The selected div may not be the content container, or extraction may occur before the page has populated it. Confirm which element matched and whether its text is present at extraction time. Use a more specific selector or wait for the content element.
Chromium is not available on the first run Pyppeteer may not yet have downloaded its browser. Run pyppeteer-install in advance, then verify the browser data directory for the same Linux user and environment that runs the script.
An evaluation expression is treated as a function rather than an expression The string passed to page.evaluate may be interpreted as a function body. For an expression such as document.body.textContent, use force_expr=True, as described in the Pyppeteer project guide: await page.evaluate("document.body.textContent", force_expr=True).
Browser processes remain after a failure The script did not close the browser when an exception occurred. Put browser cleanup in a finally block, as in the main example, so closure is attempted on both successful and failed runs.

Pyppeteer maps the JavaScript Puppeteer-style selector shortcuts to Python-friendly methods: use querySelector for $, querySelectorAll for $$, and xpath for $x. The project also documents the shorthands J, JJ, and Jx. In Python code, the named methods are usually clearer to readers who are not familiar with those aliases.

Run an extraction job predictably

  • Separate navigation from selection: navigate to the intended URL, wait for a page condition appropriate to the site, then locate the content selector.
  • Define the no-result behavior: choose whether a missing div should raise a useful application error, be retried after a wait, or count as a valid empty result.
  • Close the browser on every path: keep browser creation and cleanup in a try/finally structure. If your program creates multiple pages or browsers, give each its own clear lifetime.
  • Pin the runtime for repeatability: because the project is described as unmaintained, validate the chosen Python and Chromium combination instead of assuming updates will keep a long-lived scraper working.
  • Keep expectations realistic: pages can change their markup, move content into a client-rendered region, or block automated browsing. A selector-based script should treat a changed page as a normal failure mode and report the URL and selector that failed.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a visual capture rather than extracting a string for Python processing, ScreenshotNeo offers a screenshot API. It does not replace Pyppeteer’s text extraction: it returns an image or PDF, not the div’s text value. A GET request can capture a page without installing and managing Chromium in your own script. The cURL example below saves a WebP image; the API documentation lists the available parameters and response details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ScreenshotNeo API documentation

curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://example.com 
  -o shot.webp
  • Before capture, cookie and consent banners can be accepted and more than 60 known consent platforms, newsletter popups, and chat widgets can be removed; each of those steps can be turned off.
  • Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers identify the page verdict and whether the request was billed.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free plan and get 1,000 screenshots a month with no card.

When Pyppeteer is the right fit

Pyppeteer is useful when you need a browser to load a web page and want Python code to inspect its DOM, select an element, and return text for further processing. Use querySelectorEval for one match, querySelectorAllEval for a collection, and an element handle when you want to make the missing-element case explicit. For a new production scraper, weigh that convenience against the project’s unmaintained status and evaluate a maintained alternative before committing to a long-term runtime.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.