The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Short answer: Puppeteer is officially a JavaScript browser-automation library, not a Python library. pyppeteer is an unofficial Python port, and its maintainers say it is unmaintained. More importantly, Google says automated Search queries and scraping results without express permission violate its spam policies and Terms of Service. So there is no responsible, supported recipe here for scraping live Google Shopping results. If you own the product catalog, use Google’s supported product-data and structured-data approaches; if you need to learn browser automation, use an authorized test page instead.
Can you scrape Google Shopping with Puppeteer and Python?
Not as a supported, policy-compliant method for collecting live Google Shopping results without permission. Google identifies automated queries and scraping Search results without express permission as machine-generated traffic that violates its Search spam policies and Terms of Service. This is Google’s stated policy; this article does not make a broader legal determination.
There is also a tooling mismatch in the title. The official Puppeteer documentation describes Puppeteer as a JavaScript library. Python developers may encounter pyppeteer, but it is an unofficial port and its project repository says it is unmaintained. Neither fact establishes a stable or authorized interface for extracting Shopping listings.
The documentation reviewed for this topic does not establish current Shopping selectors, page structure, pagination behavior, result counts, or a reliable extraction technique. Google can change its page markup, and DOM-based automation depends on that markup. Publishing selectors as if they were stable would mislead you and encourage a brittle approach.
#1 Best Overall
Choose the right route for your actual goal
| Goal | Better route | Why |
|---|---|---|
| You own the products and want Google to understand or show them | Use Google’s supported product-data sharing methods and structured data | These are intended for merchant-provided product information. See Google’s ecommerce SEO guidance. |
| You want to learn browser automation | Automate a page you own or have permission to test | Puppeteer documents browser interactions such as querying elements, clicking, typing, and network request handling, without implying permission to automate any particular website. |
| You need live Shopping results from Google | Obtain express permission and confirm an approved access method before automating | Google’s published policy is the relevant boundary; a browser library does not grant access rights. |
What Puppeteer and pyppeteer actually provide
Official Puppeteer is JavaScript
Puppeteer automates Chrome and Firefox through Chrome DevTools Protocol and WebDriver BiDi. Its documented browser capabilities include locating DOM elements, clicking, typing, and intercepting or modifying network requests and responses. These are general automation features, not a Google Shopping API and not a guarantee that a result page can be extracted reliably.
pyppeteer is an unofficial, unmaintained Python port
The pyppeteer repository identifies the package as an unofficial Python port and says it is unmaintained. It lists Python 3.8 or later and installation with pip; the first run may download Chromium if a suitable browser is not installed. Those project details can change, so check the repository’s current instructions and compatibility notes before relying on it. For a new maintained Python automation workflow, evaluate a currently supported Python browser-automation library rather than assuming pyppeteer tracks official Puppeteer.
Learn the mechanics on a page you are authorized to automate
The following is a minimal Python example using pyppeteer against a local HTML file you control. It demonstrates opening a page, selecting a DOM element, and reading text. It is deliberately not a Google Shopping scraper. Save a test page as test.html in the same directory before running it:
Rank #2
<!doctype html>
<html>
<body>
<h1 class="product-title">Sample product</h1>
</body>
</html>
Install and run the script in a virtual environment. pyppeteer’s repository says Python 3.8 or later; confirm the package and browser compatibility for your environment.
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
python -m pip install pyppeteer
import asyncio
from pathlib import Path
from pyppeteer import launch
async def main():
page_path = Path("test.html").resolve().as_uri()
browser = await launch(headless=True)
try:
page = await browser.newPage()
await page.goto(page_path, {"waitUntil": "domcontentloaded"})
title = await page.querySelectorEval(
".product-title", "element => element.textContent.trim()"
)
print(title)
finally:
await browser.close()
asyncio.run(main())
Expected output is Sample product. The CSS selector is valid only because it matches the HTML you created. On a site you do not control, selectors and page behavior can change; use automation only within the authorization you have.
Use merchant-owned data instead of collecting Shopping result pages
If you are the merchant, Google’s ecommerce documentation describes supported ways to share product information and structured data so Google can understand and present your products. Start with that guidance rather than trying to reverse-engineer consumer-facing Shopping results. Google’s crawler documentation also says Storebot-Google crawl preferences affect Google’s crawling across Shopping surfaces. Those preferences concern how Google crawls a merchant’s own pages; they do not give third parties permission to scrape Google’s result pages.
- Use Google’s documented merchant product-data options to provide catalog information.
- Use structured data on product pages as described in Google’s ecommerce guidance.
- Treat crawler settings for Storebot-Google as controls over Google’s crawler behavior on your pages, not as authorization for your own scraper.
Where hosted browser automation fits
For an authorized browser workflow—such as testing your own storefront—you can run Chromium in a managed environment. Google Cloud’s Cloud Run browser automation documentation describes installing Chromium and using higher-level automation libraries such as Puppeteer or Playwright, as well as the Chrome DevTools Protocol. It lists broad browser automation use cases, including web scraping and data extraction. That general deployment capability does not override Google’s access policy for Search results.
When planning an authorized job, account for browser startup, page readiness, and the fragility of DOM-dependent selectors. Wait for a condition meaningful to your own page rather than assuming a fixed delay makes content ready. Capture logs and errors, close browser processes reliably, and test after changes to the site or browser version. Do not use hosting, request interception, or browser configuration to evade access controls.
Or skip the browser setup
For a page you are authorized to capture, ScreenshotNeo returns a screenshot or PDF from one GET request; it is not a way to bypass Google’s rules or permission requirements. Its clean-shot flow accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture, with each step configurable. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers indicate the page verdict and billing status. An MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 shots per month with no card, and paid plans start at $5 for 3,000.
Example request (replace the URL with a page you may access):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for parameters and response details. ScreenshotNeo supports PNG, JPEG, WebP, and PDF output, alongside options including full-page capture, element capture, viewport and device presets, custom CSS or JavaScript, and caching. For authorized automation that needs a captured page rather than browser setup, sign up for 1,000 free screenshots a month with no card.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting an authorized test
pyppeteer cannot find or launch Chromium
On first use, pyppeteer may download Chromium if a suitable browser is unavailable. Check the project’s current install guidance, confirm that the browser download completed, and verify that your runtime permits launching a browser process. In restricted containers, browser dependencies and launch permissions may need to be configured; do not assume a local setup transfers unchanged to a hosted runtime.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →The script cannot find an element
For your own test page, confirm that the selector exists in the rendered DOM and that the script waits until the page has loaded the relevant content. A selector copied from a different page version may no longer match. The official Puppeteer capabilities establish DOM querying generally, not any current Google Shopping selector.
Best Value
The page is blank or content is missing
Check whether navigation completed, whether the content is added after initial document loading, and whether your authorized test page needs a different readiness condition. A screenshot or DOM query taken too early can miss later content. Do not respond to a block or CAPTCHA by trying to disguise automation or defeat the control.
The workflow works locally but fails in Cloud Run
Compare the deployed Chromium installation, launch configuration, runtime permissions, and timeout limits with the local environment. Follow the Cloud Run browser automation documentation for its supported deployment pattern. A managed runtime can host a browser; it cannot grant permission to access a target site.
Cost, performance, and reliability considerations
For a controlled page, browser automation adds process startup and rendering work compared with a direct data feed. Reusing browser instances can reduce repeated startup overhead in a permitted workload, but requires careful cleanup and isolation. DOM-driven extraction also needs maintenance whenever the page markup changes. No source cited here establishes Google Shopping scraping success rates, a dependable selector set, result volumes, or a cost estimate; those should not be inferred.
Recommended Free Tools
If you own the catalog, first-party product data and structured data avoid building a scraper around presentation markup. If your authorized task specifically requires rendering a page, choose a maintained toolchain, pin and test compatible browser dependencies, and monitor for navigation failures rather than silently treating incomplete pages as complete data.
Frequently Asked Questions
Can I use official Puppeteer from Python?
Official Puppeteer is a JavaScript library. pyppeteer is an unofficial Python port, and its project repository says it is unmaintained.
Does configuring Storebot-Google allow me to scrape Google Shopping?
No. Storebot-Google settings concern Google’s crawling of your merchant pages; they do not establish third-party permission to scrape Shopping result pages.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

