Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

You may automate Glassdoor with Puppeteer only when Glassdoor has given you express written permission or you are using an approved agreement, feed, or interface. Its February 17, 2024 Terms of Use prohibit introducing automated agents to scrape, strip, or mine the service without that permission. Once access is authorized, Puppeteer can render pages, wait for content, extract narrowly defined fields, and save results without bypassing security controls.

Get authorization before writing a scraper

Authorization is a prerequisite, not a setting you add after the code works. Glassdoor’s Terms of Use revised February 17, 2024 state: “Introduce software or automated agents to the services, or access the services so as to produce multiple accounts, generate automated messages, or to scrape, strip, or mine data from the services without our express written permission.” The same terms address lawful use, account information, user content, intellectual-property rights, commercial use, and attempts to circumvent security features.

Obtain a written agreement that identifies the account, domains, URL patterns, fields, request rate, retention period, permitted purpose, and people allowed to access the output. If Glassdoor offers you a feed, partner interface, export, or other approved method, use that instead of browser automation where it supplies the required data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Keep the approval and its expiration date in your project records.
  • Define a stop condition for revoked permission, a terms change, a blocked response, or an unexpected page.
  • Collect only fields required for the approved purpose. Do not gather reviewer identities, private account data, or unrelated user content.
  • Document whether your use is internal, research, or commercial and who may receive the results.

Pick the least fragile authorized access method

Puppeteer is useful when you are authorized to read information that exists only after a browser executes JavaScript. It is not a license to collect data. Compare the available paths before choosing it:

Method When it fits Typical strengths Controls to confirm
Approved Glassdoor interface or agreement The provider supplies a documented endpoint or workflow for your use Clear authorization and defined fields Permitted fields, rate limits, retention, privacy duties, and pricing in the agreement
Licensed data feed You need recurring structured data without rendering pages Stable schema and predictable integration Freshness, redistribution rights, quotas, and total contract cost
First-party export An account owner can export the records you need Small, auditable batches with no crawler Export format, schedule, access roles, and retention
Puppeteer Your approved source requires a real browser and rendered DOM Can perform navigation, clicks, form filling, extraction, screenshots, PDFs, and tracing Selector changes, browser resource use, request volume, and permission scope
API client for an approved endpoint The authorized interface returns structured responses Less selector maintenance and lower rendering overhead Authentication, schema versioning, pagination, quotas, and error handling

Set up Puppeteer for an authorized target

Install the project

Use a current Node.js release supported by your chosen Puppeteer version. The Puppeteer security policy and Puppeteer User Guide describe installation and browser-launch options.

  1. Create a directory and initialize it: mkdir authorized-glassdoor-reader && cd authorized-glassdoor-reader && npm init -y.
  2. Install Puppeteer: npm install puppeteer.
  3. Set an environment variable to a URL covered by your written approval: export AUTHORIZED_URL='https://your-authorized-domain.example/page'.
  4. Save the script below as extract.js and run node extract.js.

The selectors in this example are deliberately generic. Replace them only with selectors documented or tested for the pages you are allowed to access; do not copy selectors from an unauthorized Glassdoor crawler.

const puppeteer = require('puppeteer');

const targetUrl = process.env.AUTHORIZED_URL;
if (!targetUrl) throw new Error('Set AUTHORIZED_URL to an approved URL');

const sleep = (ms) => new Promise(resolve => setTimeout(resolve, ms));

(async () => {
  const browser = await puppeteer.launch({
    headless: true,
    args: ['--no-sandbox', '--disable-setuid-sandbox']
  });
  const page = await browser.newPage();
  page.setDefaultNavigationTimeout(45_000);
  page.setDefaultTimeout(10_000);

  try {
    await page.setExtraHTTPHeaders({
      'Accept-Language': 'en-US,en;q=0.9'
    });

    const response = await page.goto(targetUrl, {
      waitUntil: 'domcontentloaded',
      timeout: 45_000
    });
    if (!response || !response.ok()) {
      throw new Error(`Navigation failed: ${response ? response.status() : 'no response'}`);
    }

    // Wait for a selector that your authorization and page contract specify.
    const requiredSelector = process.env.REQUIRED_SELECTOR || 'main';
    await page.waitForSelector(requiredSelector, { visible: true });
    await sleep(500); // Allow a permitted client-side render to settle.

    const records = await page.$$eval('[data-authorized-record]', nodes =>
      nodes.map(node => ({
        title: node.querySelector('[data-title]')?.textContent?.trim() || null,
        location: node.querySelector('[data-location]')?.textContent?.trim() || null,
        rating: node.querySelector('[data-rating]')?.textContent?.trim() || null
      }))
    );

    const missing = records.filter(row => !row.title);
    if (missing.length) {
      throw new Error(`Validation failed: ${missing.length} records lack title`);
    }

    console.log(JSON.stringify({
      url: page.url(),
      capturedAt: new Date().toISOString(),
      count: records.length,
      records
    }, null, 2));
  } finally {
    await browser.close();
  }
})().catch(error => {
  console.error(error.stack || error);
  process.exitCode = 1;
});

This workflow launches Chromium, navigates once, waits for a known page state, selects only the approved fields, validates a required field, and closes the browser in a finally block. If your approved page uses pagination, add one page at a time with a documented maximum; do not recursively follow every link.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make extraction narrow, private, and auditable

Use an explicit field contract

Write down each field, its selector or source, its type, and whether it may be retained. Treat a changed selector or a missing field as a validation failure rather than silently storing a different value. Normalize whitespace and dates only after preserving the source value when your agreement requires provenance.

Protect account and user data

Use a dedicated account with the minimum permissions needed. Keep credentials in environment variables or a secret manager, never in source control or page logs. Redact cookies, authorization headers, reviewer names, email addresses, and free-text content unless the written scope expressly includes them. Encrypt stored output and set an automatic deletion date.

Throttle and cache

Queue URLs at a low, agreed rate. Cache successful responses using a key that includes the URL and relevant parameters, and avoid revisiting unchanged pages. A retry should use exponential backoff with jitter and a small maximum attempt count; it must not turn a denial into a flood of requests.

Log decisions, not sensitive payloads

Record timestamp, URL pattern, response status, duration, selector version, retry count, and the authorization reference. Do not put page HTML or personal data in ordinary logs. Keep an audit trail for exports and deletions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not defeat bot checks or security controls

Automation defenses commonly inspect signals such as navigator.webdriver and JavaScript challenges, as described in the EURECOM research on automation signals and anti-bot challenges. A CAPTCHA, interstitial, unusual redirect, or “access denied” page is a signal to stop and contact the data owner—not an invitation to install a stealth plugin, rotate proxies, spoof fingerprints, solve CAPTCHAs, or evade rate limits. Attempts to circumvent security features can violate the site’s restrictions and your agreement.

Detect these states explicitly. Save only a minimal diagnostic (status, title, and timestamp), stop the queue, and ask the authorized contact whether a supported interface or allowlist is available. Never submit credentials or automated messages to an account you do not control.

Or skip the browser setup

If you need an authorized visual capture rather than structured Glassdoor records, ScreenshotNeo returns a PNG, JPEG, WebP, or PDF from one GET request. It is a screenshot API, not a way to bypass Glassdoor permission requirements or extract reviewer data. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and whether it was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Use a URL covered by your authorization and read the ScreenshotNeo documentation for options:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/authorized-page -o shot.webp

There is a free plan with 1,000 screenshots per month and no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try an authorized capture.

Performance, reliability, and cost controls

  • Browser lifetime: Reuse one browser for a small batch but create a fresh page per URL. Close pages and the browser on every success and failure to prevent memory leaks.
  • Waiting: Prefer a specific selector or documented application-ready signal over a long fixed delay. Use network-idle waits only when the authorized page has predictable background traffic.
  • Concurrency: Start with one page at a time. Increase concurrency only if the agreement permits it and monitoring shows stable memory, response times, and error rates.
  • Retries: Retry transient network errors with capped exponential backoff. Do not retry a policy denial, CAPTCHA, authentication failure, or selector-validation failure automatically.
  • Cost: Puppeteer consumes your compute, bandwidth, storage, and engineering time. Measure average browser startup time, page duration, memory per page, and cache-hit rate before setting a budget.
  • Change management: Pin and review the Puppeteer version, test selectors against fixtures, and alert when record counts suddenly drop or required fields become null.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

Symptom Likely cause Compliant fix
TimeoutError waiting for a selector The page changed, content is not in scope, or the selector is wrong Confirm the approved URL and selector contract, capture a minimal DOM diagnostic, then update the selector under change control.
Navigation returns 401 or 403 Missing permission, expired session, or blocked access Stop retries; verify authorization and credentials with the provider.
A CAPTCHA or bot challenge appears The service detected automation Do not bypass it. Stop the job and request an approved feed, allowlist, or alternate workflow.
Records are empty but the page looks loaded Data is rendered later, inside a frame, or behind an interaction Wait for the documented ready selector, inspect permitted frames, and perform only approved clicks. Revalidate fields before storing anything.
Chromium fails in a container Missing OS libraries, sandbox restrictions, or insufficient shared memory Install dependencies using your deployment image, follow Puppeteer’s security guidance, and fix the runtime rather than changing site-detection behavior.
Memory grows during a batch Pages, listeners, or browser processes are not closed Close each page in a finally block, cap batch size, monitor RSS, and restart the browser at a planned boundary.

FAQ

Can I scrape public Glassdoor pages without logging in?

Public visibility does not remove the Terms of Use restriction on automated scraping. You still need express written permission or an approved interface.

Should I store the HTML for later reprocessing?

Only if your authorization expressly allows it and your retention policy requires it. Otherwise retain the smallest structured result and a non-sensitive audit record.

Is Puppeteer preferable to an approved API?

Use the API when it supplies the fields you need. Choose Puppeteer only when authorized data requires browser rendering and you can maintain selectors, browser resources, and privacy controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I scrape public Glassdoor pages without logging in?

Public visibility does not remove the Terms of Use restriction on automated scraping. You still need express written permission or an approved interface.

Should I store the HTML for later reprocessing?

Only if your authorization expressly allows it and your retention policy requires it. Otherwise retain the smallest structured result and a non-sensitive audit record.

Is Puppeteer preferable to an approved API?

Use the API when it supplies the fields you need. Choose Puppeteer only when authorized data requires browser rendering and you can maintain selectors, browser resources, and privacy controls.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.