Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The usual reason is that the data is not in the page’s initial HTTP response. A traditional scraper downloads that response and parses its HTML. A modern app may return only a small shell, then run JavaScript that calls an API, GraphQL endpoint, or other XHR/fetch request and inserts the results into the live page. Your browser executes those steps; a basic requests, curl, or Scrapy request does not.

That is why View Source can look empty while Inspect Element shows a populated table. The reliable fix is to identify the request carrying the data and reproduce it directly when possible. If JavaScript execution, clicks, login state, scrolling, or browser-only storage is required, use a browser context such as Playwright. The workflow below shows how to decide, implement, wait correctly, diagnose failures, and stay within the site’s authorization and usage rules.

View Source and Inspect Element are different documents

View Source is the server’s original response. Inspect Element displays the current DOM after scripts have modified it. JavaScript can create elements, replace placeholders, paginate results, and hydrate an app-shell page after the response has arrived. Google Search Central describes this app-shell pattern: some JavaScript sites send initial HTML without the actual content and require JavaScript execution before the generated content exists.

What you inspect What it contains What it tells you
HTTP response / View Source Markup and data returned by the server for that request Whether a plain HTTP scraper can see the data immediately
Live DOM / Inspect Element The document after JavaScript, user actions, and asynchronous requests What a user sees after the page has run
Network panel Document, XHR/fetch, JSON, GraphQL, image, and other requests Which request actually delivered the missing data

If the value is absent from the saved response, changing a CSS selector will not help. You must either call the data endpoint or run the page in a browser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A diagnostic workflow that finds the missing step

  1. Save the scraper’s exact response

    Record the URL, method, query string, request body, status code, response headers, and body returned by your HTTP client. Compare that body with View Source, not with Inspect Element. If the value is missing in both, the problem is not your selector; it is a later request or rendering step.

  2. Watch the request that reveals the data

    Open developer tools, select Network, filter for Fetch/XHR, JSON, or GraphQL, and reload the page. Then perform the action that reveals the data: submit a search, change a filter, click “Next,” or scroll. Inspect response previews and payloads until you find the rows or fields you need.

  3. Reproduce the data request directly

    Use the request’s method, URL, query parameters or JSON body, required headers, cookies, and authentication token. This is usually faster and lighter than rendering a full page. Scrapy’s guidance is to download the page with an HTTP client first and use browser network tools to locate the follow-up request when the desired data is not in the response.

  4. Check whether the request is stable

    Replay it outside the browser and inspect the result. Some applications issue short-lived tokens, sign parameters, rotate cookies, or require a preliminary request. If the direct call works repeatedly with permitted credentials, build your scraper around that endpoint. If it only works after JavaScript creates state or after a user interaction, move to browser automation.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  5. Choose a meaningful readiness condition

    A load event only means the document’s loading phase completed. Wait for the specific response, a selector that contains real data, or a documented network-idle condition. Avoid an arbitrary short sleep: it may pass on a fast run and fail under normal latency.

Choose the lightest approach that can see the data

Approach Use it when Advantages Costs and risks
Direct HTTP/API reproduction The endpoint is visible, stable, and accessible with permitted credentials Fast, low CPU, structured JSON, easy to scale You must maintain parameters, tokens, cookies, and request changes
Playwright or equivalent browser JavaScript, clicks, scrolling, iframes, storage, or a complex login flow is required Closest to the user-visible behavior; can observe requests and DOM changes More CPU, startup time, browser maintenance, and synchronization work
Managed browser rendering You want hosted execution and rendered output without operating browsers Offloads browser infrastructure and can provide rendered pages or element captures Service limits, per-use pricing, and less control than your own context

Reproduce the endpoint without a browser

Start with a request captured from your own authorized session. Keep the method, body, and headers exact; do not assume a visible page URL is the data URL.

cURL

curl --fail-with-body -sS -X POST "$API_URL" 
  -H 'Accept: application/json' 
  -H "Authorization: Bearer $API_TOKEN" 
  -H 'Content-Type: application/json' 
  --data "$API_BODY"

Set API_URL, API_TOKEN, and API_BODY from a request you are allowed to make. For a GET request, use -G and add the captured query parameters with -d.

Python with requests

import json
import os
import requests

url = os.environ['API_URL']
token = os.environ.get('API_TOKEN')
body = json.loads(os.environ.get('API_BODY', '{}'))
headers = {'Accept': 'application/json'}
if token:
    headers['Authorization'] = f'Bearer {token}'

response = requests.post(url, headers=headers, json=body, timeout=30)
response.raise_for_status()
data = response.json()
for row in data.get('items', data if isinstance(data, list) else []):
    print(row)

Preserve pagination fields such as cursors or page numbers, and stop when the server indicates there are no more results. Do not log access tokens or personal data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Playwright when execution or interaction is required

Playwright contexts run JavaScript and can carry cookies, HTTP credentials, proxies, and other browser state. Its network APIs let you wait for the response triggered by a click instead of guessing how long the page needs.

Python example: capture a response after a search

import asyncio
import json
import os
from playwright.async_api import async_playwright

async def main():
    target = os.environ['TARGET_URL']
    search_term = os.environ.get('SEARCH_TERM', 'example')

    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        context = await browser.new_context(
            java_script_enabled=True,
            storage_state=os.environ.get('STORAGE_STATE') or None,
        )
        page = await context.new_page()
        await page.goto(target, wait_until='domcontentloaded')
        await page.get_by_role('textbox').fill(search_term)
        async with page.expect_response(
            lambda r: '/api/' in r.url and r.request.method in ('GET', 'POST')
        ) as response_info:
            await page.get_by_role('button', name='Search').click()
        response = await response_info.value
        if not response.ok:
            raise RuntimeError(f'{response.status} from {response.url}')
        payload = await response.json()
        print(json.dumps(payload, ensure_ascii=False))
        await browser.close()

asyncio.run(main())

Install the package and browser once with pip install playwright followed by playwright install chromium. Replace the role-based locators and /api/ predicate with values observed in your site’s Network panel. If the application loads data on initial navigation, wrap page.goto in page.expect_response instead.

Extract rendered text only when no usable endpoint exists

If the response is encrypted, assembled from several calls, or otherwise impractical to reproduce, wait for a data-bearing selector and read the DOM:

await page.locator('[data-testid="results"] li').first.wait_for(state='visible')
rows = await page.locator('[data-testid="results"] li').all_inner_texts()

Prefer the response payload when it contains the authoritative fields. DOM text can include formatting, hidden labels, duplicates, or localized values.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle browser state, frames, and network interception

Authentication and cookies

A page may show data only after a login cookie, CSRF token, HTTP credential, or custom header is present. Use a permitted account and a stored Playwright authentication state, or reproduce the documented login flow. Never defeat an access control or use credentials you do not have permission to use.

iframes

Data inside an iframe belongs to that frame’s document. Locate the frame by URL or title and query it separately. A cross-origin frame may prevent DOM access even though the browser displays it; in that case, capture the frame’s own permitted network request or use the provider’s API.

Service workers

Service workers can intercept requests and make them absent from ordinary routing events. If network events appear incomplete, test a context with service workers blocked, then verify that the page still behaves correctly:

context = await browser.new_context(service_workers='block')

Blocking them is a diagnostic option, not a universal fix. Some applications depend on a service worker for their data path.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CORS and opaque responses

Cross-origin permissions govern whether page JavaScript can read a response. A no-cors fetch produces an opaque response that browser JavaScript cannot inspect. CORS is enforced by browsers; a server-side scraper is not subject to the page’s browser origin check, but it still needs authorization and the correct request credentials.

Waiting, pagination, and lazy content

  • Wait for a response when a click or filter triggers a known endpoint.
  • Wait for a selector containing real content, not merely an empty table element.
  • Use network-idle carefully. Analytics, ads, and long-polling can prevent true idleness. A response or content condition is usually more precise.
  • Scroll deliberately for infinite lists, and stop when the item count or cursor stops changing.
  • Capture each pagination cursor and deduplicate records by a stable identifier.

Lazy-loaded images and rows may not exist until they approach the viewport. If your goal is structured data, trigger the same scroll or pagination action a user would and collect the resulting payloads.

Common failures and fixes

Symptom Likely cause Fix
Empty HTML but populated browser page App shell or client-side rendering Find the XHR/fetch response; otherwise run a JavaScript-enabled browser.
Selector returns zero elements Selector ran before rendering, or the data is in an iframe Wait for a data-bearing selector and inspect frame boundaries.
Direct endpoint returns 401 or 403 Missing, expired, or unauthorized credentials Refresh the permitted session, include required cookies or headers, and verify scope.
Response is HTML instead of JSON Redirect, login page, bot check, or wrong endpoint Log final URL and status, disable automatic assumptions, and inspect the response body.
Works manually but times out in automation Arbitrary sleep, slow dependency, or a never-idle analytics request Wait for the specific response or selector and set realistic navigation and action timeouts.
Network listener sees nothing Service worker interception or listener attached too late Attach the listener before navigation and test service_workers='block'.
Rows differ between runs Personalization, locale, timezone, rotating tokens, or changing backend data Set the intended locale/timezone, persist authorized state, record request parameters, and store retrieval timestamps.
Browser sees a challenge page Bot protection or an access policy Do not attempt to bypass it. Obtain permission, use an official API, or ask the site owner for an approved integration.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your deliverable is a faithful screenshot or PDF rather than structured rows, ScreenshotNeo runs the capture for you. Before the shot it accepts the cookie or consent banner and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the result in X-Page-Verdict and X-Billed headers.

Use the API documentation at screenshotneo.com/docs/ for options such as full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper size and page ranges, custom CSS or JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification. It also accepts the parameter names used by other screenshot APIs, which can simplify migration.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo is not a substitute for an authorized JSON endpoint when you need records for analysis. It is useful when the required output is visual evidence, a rendered page, or a PDF, and its MCP server gives Claude, Cursor, and other MCP clients take_screenshot, get_page_info, and capture_pdf tools.

Plan Included shots Price
Free 1,000 per month $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Every feature is available on every plan, and yearly billing gives two months free. Start with 1,000 free screenshots a month without a card.

Performance, reliability, and operating costs

Direct API calls normally win on latency and throughput because they avoid browser startup, layout, fonts, images, and JavaScript execution. Browser automation is heavier, so reuse a browser process, create isolated contexts, cap concurrency, and block unnecessary resources only when doing so does not alter the data path. Cache immutable responses and honor server rate limits.

For reliability, record the request URL, status, response schema, page URL, locale, and timestamp. Validate required fields and alert on schema changes instead of silently writing empty records. Retry transient network failures with bounded backoff, but do not aggressively retry authorization failures or policy blocks. Keep a small fixture of known responses for regression tests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Managed rendering trades infrastructure work for a service request and its plan limits. ScreenshotNeo’s billing behavior is explicit through its verdict and billed headers, which helps distinguish a failed or cached capture from a chargeable clean shot.

Respect authorization and site policy

Visible data is not automatically unrestricted data. Check robots directives, terms, rate limits, privacy obligations, and account permissions. Use official APIs where available, identify your crawler when appropriate, minimize personal data, and retain only what you need. Authentication barriers should be handled with valid credentials and the owner’s permission, never by defeating access controls.

FAQ

Why does “Ctrl+F” find text that my downloaded HTML does not?

The browser’s find operation searches the live DOM after scripts have inserted content. Search the saved HTTP response separately to confirm whether the text was ever delivered initially.

Can changing the User-Agent make the missing data appear?

It can change which variant a server returns, but it does not execute JavaScript. Treat it as a diagnostic variable, not a replacement for locating the data request or running a browser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a screenshot enough to recover the underlying table?

No. Optical character recognition can approximate visible text, but it loses structure and hidden fields. Use the network payload or rendered DOM for dependable structured extraction.

When should I ask the site owner for an integration?

Ask when the endpoint is private, unstable, rate-limited, protected by a challenge, or contains personal information you are not clearly authorized to collect. An approved API or export is safer than reverse-engineering a private flow.

Frequently Asked Questions

Why does “Ctrl+F” find text that my downloaded HTML does not?

The browser searches the live DOM after JavaScript inserts content; the saved response contains only the original server HTML.

Can changing the User-Agent make missing data appear?

It may change the server’s response variant, but it does not execute JavaScript. You still need the data request or a browser context.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a screenshot enough to recover an underlying table?

No. Use the network payload or rendered DOM for reliable structured extraction; OCR from an image can lose fields and structure.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.