The usual reason is that the data is not in the page’s initial HTTP response. A traditional scraper downloads that response and parses its HTML. A modern app may return only a small shell, then run JavaScript that calls an API, GraphQL endpoint, or other XHR/fetch request and inserts the results into the live page. Your browser executes those steps; a basic requests, curl, or Scrapy request does not.
That is why View Source can look empty while Inspect Element shows a populated table. The reliable fix is to identify the request carrying the data and reproduce it directly when possible. If JavaScript execution, clicks, login state, scrolling, or browser-only storage is required, use a browser context such as Playwright. The workflow below shows how to decide, implement, wait correctly, diagnose failures, and stay within the site’s authorization and usage rules.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
The Proxy Playbook: The Complete Guide to Proxy Servers: How to Source, Test, and Scale Residential,... | $29.95 | Buy on Amazon |
| 2 |
|
How to Host your own Web Server | $15.60 | Buy on Amazon |
View Source and Inspect Element are different documents
View Source is the server’s original response. Inspect Element displays the current DOM after scripts have modified it. JavaScript can create elements, replace placeholders, paginate results, and hydrate an app-shell page after the response has arrived. Google Search Central describes this app-shell pattern: some JavaScript sites send initial HTML without the actual content and require JavaScript execution before the generated content exists.
| What you inspect | What it contains | What it tells you |
|---|---|---|
| HTTP response / View Source | Markup and data returned by the server for that request | Whether a plain HTTP scraper can see the data immediately |
| Live DOM / Inspect Element | The document after JavaScript, user actions, and asynchronous requests | What a user sees after the page has run |
| Network panel | Document, XHR/fetch, JSON, GraphQL, image, and other requests | Which request actually delivered the missing data |
If the value is absent from the saved response, changing a CSS selector will not help. You must either call the data endpoint or run the page in a browser.
#1 Best Overall
A diagnostic workflow that finds the missing step
-
Save the scraper’s exact response
Record the URL, method, query string, request body, status code, response headers, and body returned by your HTTP client. Compare that body with View Source, not with Inspect Element. If the value is missing in both, the problem is not your selector; it is a later request or rendering step.
-
Watch the request that reveals the data
Open developer tools, select Network, filter for
Fetch/XHR, JSON, or GraphQL, and reload the page. Then perform the action that reveals the data: submit a search, change a filter, click “Next,” or scroll. Inspect response previews and payloads until you find the rows or fields you need. -
Reproduce the data request directly
Use the request’s method, URL, query parameters or JSON body, required headers, cookies, and authentication token. This is usually faster and lighter than rendering a full page. Scrapy’s guidance is to download the page with an HTTP client first and use browser network tools to locate the follow-up request when the desired data is not in the response.
-
Check whether the request is stable
Replay it outside the browser and inspect the result. Some applications issue short-lived tokens, sign parameters, rotate cookies, or require a preliminary request. If the direct call works repeatedly with permitted credentials, build your scraper around that endpoint. If it only works after JavaScript creates state or after a user interaction, move to browser automation.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy. -
Choose a meaningful readiness condition
A
loadevent only means the document’s loading phase completed. Wait for the specific response, a selector that contains real data, or a documented network-idle condition. Avoid an arbitrary short sleep: it may pass on a fast run and fail under normal latency.
Choose the lightest approach that can see the data
| Approach | Use it when | Advantages | Costs and risks |
|---|---|---|---|
| Direct HTTP/API reproduction | The endpoint is visible, stable, and accessible with permitted credentials | Fast, low CPU, structured JSON, easy to scale | You must maintain parameters, tokens, cookies, and request changes |
| Playwright or equivalent browser | JavaScript, clicks, scrolling, iframes, storage, or a complex login flow is required | Closest to the user-visible behavior; can observe requests and DOM changes | More CPU, startup time, browser maintenance, and synchronization work |
| Managed browser rendering | You want hosted execution and rendered output without operating browsers | Offloads browser infrastructure and can provide rendered pages or element captures | Service limits, per-use pricing, and less control than your own context |
Reproduce the endpoint without a browser
Start with a request captured from your own authorized session. Keep the method, body, and headers exact; do not assume a visible page URL is the data URL.
cURL
curl --fail-with-body -sS -X POST "$API_URL"
-H 'Accept: application/json'
-H "Authorization: Bearer $API_TOKEN"
-H 'Content-Type: application/json'
--data "$API_BODY"
Set API_URL, API_TOKEN, and API_BODY from a request you are allowed to make. For a GET request, use -G and add the captured query parameters with -d.
Python with requests
import json
import os
import requests
url = os.environ['API_URL']
token = os.environ.get('API_TOKEN')
body = json.loads(os.environ.get('API_BODY', '{}'))
headers = {'Accept': 'application/json'}
if token:
headers['Authorization'] = f'Bearer {token}'
response = requests.post(url, headers=headers, json=body, timeout=30)
response.raise_for_status()
data = response.json()
for row in data.get('items', data if isinstance(data, list) else []):
print(row)
Preserve pagination fields such as cursors or page numbers, and stop when the server indicates there are no more results. Do not log access tokens or personal data.
Use Playwright when execution or interaction is required
Playwright contexts run JavaScript and can carry cookies, HTTP credentials, proxies, and other browser state. Its network APIs let you wait for the response triggered by a click instead of guessing how long the page needs.
Python example: capture a response after a search
import asyncio
import json
import os
from playwright.async_api import async_playwright
async def main():
target = os.environ['TARGET_URL']
search_term = os.environ.get('SEARCH_TERM', 'example')
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context(
java_script_enabled=True,
storage_state=os.environ.get('STORAGE_STATE') or None,
)
page = await context.new_page()
await page.goto(target, wait_until='domcontentloaded')
await page.get_by_role('textbox').fill(search_term)
async with page.expect_response(
lambda r: '/api/' in r.url and r.request.method in ('GET', 'POST')
) as response_info:
await page.get_by_role('button', name='Search').click()
response = await response_info.value
if not response.ok:
raise RuntimeError(f'{response.status} from {response.url}')
payload = await response.json()
print(json.dumps(payload, ensure_ascii=False))
await browser.close()
asyncio.run(main())
Install the package and browser once with pip install playwright followed by playwright install chromium. Replace the role-based locators and /api/ predicate with values observed in your site’s Network panel. If the application loads data on initial navigation, wrap page.goto in page.expect_response instead.
Extract rendered text only when no usable endpoint exists
If the response is encrypted, assembled from several calls, or otherwise impractical to reproduce, wait for a data-bearing selector and read the DOM:
await page.locator('[data-testid="results"] li').first.wait_for(state='visible')
rows = await page.locator('[data-testid="results"] li').all_inner_texts()
Prefer the response payload when it contains the authoritative fields. DOM text can include formatting, hidden labels, duplicates, or localized values.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Handle browser state, frames, and network interception
Authentication and cookies
A page may show data only after a login cookie, CSRF token, HTTP credential, or custom header is present. Use a permitted account and a stored Playwright authentication state, or reproduce the documented login flow. Never defeat an access control or use credentials you do not have permission to use.
iframes
Data inside an iframe belongs to that frame’s document. Locate the frame by URL or title and query it separately. A cross-origin frame may prevent DOM access even though the browser displays it; in that case, capture the frame’s own permitted network request or use the provider’s API.
Service workers
Service workers can intercept requests and make them absent from ordinary routing events. If network events appear incomplete, test a context with service workers blocked, then verify that the page still behaves correctly:
context = await browser.new_context(service_workers='block')
Blocking them is a diagnostic option, not a universal fix. Some applications depend on a service worker for their data path.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
CORS and opaque responses
Cross-origin permissions govern whether page JavaScript can read a response. A no-cors fetch produces an opaque response that browser JavaScript cannot inspect. CORS is enforced by browsers; a server-side scraper is not subject to the page’s browser origin check, but it still needs authorization and the correct request credentials.
Waiting, pagination, and lazy content
- Wait for a response when a click or filter triggers a known endpoint.
- Wait for a selector containing real content, not merely an empty table element.
- Use network-idle carefully. Analytics, ads, and long-polling can prevent true idleness. A response or content condition is usually more precise.
- Scroll deliberately for infinite lists, and stop when the item count or cursor stops changing.
- Capture each pagination cursor and deduplicate records by a stable identifier.
Lazy-loaded images and rows may not exist until they approach the viewport. If your goal is structured data, trigger the same scroll or pagination action a user would and collect the resulting payloads.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Empty HTML but populated browser page | App shell or client-side rendering | Find the XHR/fetch response; otherwise run a JavaScript-enabled browser. |
| Selector returns zero elements | Selector ran before rendering, or the data is in an iframe | Wait for a data-bearing selector and inspect frame boundaries. |
| Direct endpoint returns 401 or 403 | Missing, expired, or unauthorized credentials | Refresh the permitted session, include required cookies or headers, and verify scope. |
| Response is HTML instead of JSON | Redirect, login page, bot check, or wrong endpoint | Log final URL and status, disable automatic assumptions, and inspect the response body. |
| Works manually but times out in automation | Arbitrary sleep, slow dependency, or a never-idle analytics request | Wait for the specific response or selector and set realistic navigation and action timeouts. |
| Network listener sees nothing | Service worker interception or listener attached too late | Attach the listener before navigation and test service_workers='block'. |
| Rows differ between runs | Personalization, locale, timezone, rotating tokens, or changing backend data | Set the intended locale/timezone, persist authorized state, record request parameters, and store retrieval timestamps. |
| Browser sees a challenge page | Bot protection or an access policy | Do not attempt to bypass it. Obtain permission, use an official API, or ask the site owner for an approved integration. |
Or skip the browser setup
If your deliverable is a faithful screenshot or PDF rather than structured rows, ScreenshotNeo runs the capture for you. Before the shot it accepts the cookie or consent banner and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the result in X-Page-Verdict and X-Billed headers.
Use the API documentation at screenshotneo.com/docs/ for options such as full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper size and page ranges, custom CSS or JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification. It also accepts the parameter names used by other screenshot APIs, which can simplify migration.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo is not a substitute for an authorized JSON endpoint when you need records for analysis. It is useful when the required output is visual evidence, a rendered page, or a PDF, and its MCP server gives Claude, Cursor, and other MCP clients take_screenshot, get_page_info, and capture_pdf tools.
| Plan | Included shots | Price |
|---|---|---|
| Free | 1,000 per month | $0, no card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Every feature is available on every plan, and yearly billing gives two months free. Start with 1,000 free screenshots a month without a card.
Performance, reliability, and operating costs
Direct API calls normally win on latency and throughput because they avoid browser startup, layout, fonts, images, and JavaScript execution. Browser automation is heavier, so reuse a browser process, create isolated contexts, cap concurrency, and block unnecessary resources only when doing so does not alter the data path. Cache immutable responses and honor server rate limits.
For reliability, record the request URL, status, response schema, page URL, locale, and timestamp. Validate required fields and alert on schema changes instead of silently writing empty records. Retry transient network failures with bounded backoff, but do not aggressively retry authorization failures or policy blocks. Keep a small fixture of known responses for regression tests.
Managed rendering trades infrastructure work for a service request and its plan limits. ScreenshotNeo’s billing behavior is explicit through its verdict and billed headers, which helps distinguish a failed or cached capture from a chargeable clean shot.
Respect authorization and site policy
Visible data is not automatically unrestricted data. Check robots directives, terms, rate limits, privacy obligations, and account permissions. Use official APIs where available, identify your crawler when appropriate, minimize personal data, and retain only what you need. Authentication barriers should be handled with valid credentials and the owner’s permission, never by defeating access controls.
FAQ
Why does “Ctrl+F” find text that my downloaded HTML does not?
The browser’s find operation searches the live DOM after scripts have inserted content. Search the saved HTTP response separately to confirm whether the text was ever delivered initially.
Can changing the User-Agent make the missing data appear?
It can change which variant a server returns, but it does not execute JavaScript. Treat it as a diagnostic variable, not a replacement for locating the data request or running a browser.
Is a screenshot enough to recover the underlying table?
No. Optical character recognition can approximate visible text, but it loses structure and hidden fields. Use the network payload or rendered DOM for dependable structured extraction.
When should I ask the site owner for an integration?
Ask when the endpoint is private, unstable, rate-limited, protected by a challenge, or contains personal information you are not clearly authorized to collect. An approved API or export is safer than reverse-engineering a private flow.
Frequently Asked Questions
Why does “Ctrl+F” find text that my downloaded HTML does not?
The browser searches the live DOM after JavaScript inserts content; the saved response contains only the original server HTML.
Can changing the User-Agent make missing data appear?
It may change the server’s response variant, but it does not execute JavaScript. You still need the data request or a browser context.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsIs a screenshot enough to recover an underlying table?
No. Use the network payload or rendered DOM for reliable structured extraction; OCR from an image can lose fields and structure.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

