Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An MCP server can make web retrieval callable by an AI application, but it does not make scraped pages trustworthy. Treat MCP as a control interface: the client selects a narrowly defined tool, the server performs a permitted fetch or browser action, and the returned page text, screenshots and metadata remain untrusted data. Secure deployments add destination restrictions, least-privilege credentials, process isolation, validation and human approval for consequential actions.

What an MCP scraping server actually is

The Model Context Protocol (MCP) standardizes how an AI application discovers and calls server-exposed tools and data capabilities. It is not a scraping engine, a browser, a malware filter or a trust certification. A server may use direct HTTP, a browser, an API client or another retrieval method internally; the MCP client sees the tools and receives structured results.

The control layer

The client sends a structured request such as “retrieve this allowed page” or “inspect the current browser tab.” The server decides whether that operation is within its policy, executes it and returns a result. Keep these operations small and explicit: fetch_article_text, list_links and capture_product_page are safer than a generic run_browser_command.

The data layer

Everything returned from a site is data to analyze, not instructions to obey. That includes visible text, hidden text, HTML attributes, JSON embedded in a page, screenshots, HTTP headers and error messages. A page can contain prompt-injection text that attempts to redirect the agent, disclose credentials or invoke another tool.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A normal request path

  1. The AI client negotiates capabilities and learns the server’s tool schemas.
  2. The client sends validated arguments for one tool.
  3. The server checks the destination, authorization context and action scope.
  4. The retrieval component performs an HTTP request or browser operation.
  5. The server labels and returns the result as untrusted external content.
  6. The client decides what to do next under the user’s original instructions.

MCP request handling is stateless at the protocol level. If pagination, a login session or a browser tab must persist across calls, the application must create an explicit, access-controlled identifier and define its lifetime.

Why “carry control, not data” matters

The phrase is an architectural rule, not a guarantee made by the MCP specification. A tool call carries authority: it can navigate, read a file, use a credential or submit a form. A page result carries information that may be false, malicious or simply irrelevant. Mixing those roles lets content silently expand the agent’s authority.

Demarcate every external result

Return a typed envelope that separates fields such as source_url, retrieved_at, content and warnings. In the client prompt, clearly mark the content as untrusted. Do not allow text inside content to override system or user instructions, authorize a new destination or approve a state-changing action.

Validate tool metadata too

Tool descriptions and schemas are supplied by the server. The MCP specification says clients and servers should document which schema dialects they support, and server identity metadata is self-reported. Treat a newly added or changed tool as a configuration change requiring review, not as proof of safety. OWASP’s MCP guidance calls out tool poisoning, “rug pulls,” cross-server influence and over-scoped tokens as risks.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep origins narrow

Give the agent an allowlist of the sites required for the task. Validate the URL before the request, after every redirect and after DNS resolution. Reject dangerous schemes such as file:, gopher: and javascript:; block private, loopback, link-local and cloud-metadata address ranges where the fetch path could reach them. The MCP security guidance discusses SSRF in OAuth metadata discovery; the same destination discipline is appropriate for scraping fetches.

How a browser-control MCP server works

Some pages require JavaScript execution, scrolling, clicks or an authenticated browser profile. In that design, the MCP server exposes bounded browser actions and returns snapshots or extracted values. Microsoft’s documented Chrome DevTools MCP example uses Puppeteer to control a Chromium-based browser, Edge or WebView2. That is one implementation example, not a requirement of MCP.

Separate browser powers into tools

  • Navigation: accept an allowlisted URL and a limited wait policy.
  • Inspection: return selected text, accessibility data or a screenshot.
  • Interaction: click or type only when the user has requested that specific action.
  • Submission: isolate form submission, publishing and account changes behind an explicit confirmation step.

Do not expose arbitrary JavaScript evaluation or an unrestricted DevTools endpoint to an agent unless the process is isolated and the risk is understood. A read-only extraction tool should not inherit permissions that allow account changes.

Security controls for a production scraping server

Constrain destinations and redirects

Parse URLs with a real URL library, require HTTP(S), normalize hostnames and enforce an origin allowlist. Re-check every redirect rather than trusting the initial URL. Resolve DNS and reject addresses in private or reserved ranges; enforce egress rules at the network layer as a second control. Set maximum response sizes, redirect counts and download times so a permitted origin cannot become an unbounded proxy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use narrowly scoped credentials

Never forward a broad bearer token or an entire browser profile to arbitrary destinations. Store credentials outside prompts, inject only the credential needed for the approved origin and redact secrets from logs and returned text. Separate read-only retrieval credentials from identities that can post, purchase, delete or modify data.

Assume local stdio is privileged

With stdio transport, the client starts the MCP server as a local subprocess. The MCP project’s Security Policy states: “Deployments that run stdio servers at reduced privilege (containers, sandboxes) are responsible for enforcing isolation at that boundary; the SDK’s stdio transport is not a sandbox.” Run the process with a dedicated account or container, a read-only filesystem where possible, no unnecessary environment variables and a restricted network policy.

Require approval for consequential actions

Make extraction and mutation different tools. Before a click that submits a form or changes an account, show the destination, the intended action and the relevant parameters, then obtain user confirmation. Do not let instructions found in a page silently satisfy that confirmation.

Log and review

Record the server and tool version, destination, authorization context, request identifier, result classification and whether state changed. Review source provenance, dependency updates, declared permissions and release notes. Keep sensitive page content out of routine logs, and define retention and deletion rules for authenticated data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choosing an implementation

There is no tested universal ranking of MCP scraping servers. Compare candidates against the risk and rendering needs of your workload:

Decision axis Questions to ask
Retrieval method Does it use direct HTTP for simple pages, a browser for rendered or interactive pages, or both?
Destination control Are origin allowlists, redirect checks, DNS validation and egress restrictions enforced server-side?
Data handling What leaves the browser, how long is it retained, and are cookies, tokens and page text redacted?
Permission model Are read-only tools separate from clicks, submissions and other state-changing actions?
Isolation What filesystem, network and operating-system privileges does the process receive?
Maintenance and provenance Is source code available for review, how are dependencies updated, and are security practices documented?

Choose the least powerful implementation that meets the page requirements. Direct HTTP is usually simpler for static, public endpoints; browser automation is justified when rendering, interaction or a real session is necessary.

DIY browser retrieval before you wrap it as an MCP tool

The following Node.js example demonstrates the browser portion only. It is not an MCP server: an MCP adapter would expose this operation through a narrow tool schema, enforce policy before calling it and return the result as untrusted data.

Prerequisites

  • Node.js 18 or newer.
  • A project with Puppeteer installed: npm install puppeteer.
  • A URL you are authorized to retrieve.
import puppeteer from 'puppeteer';

const target = process.argv[2];
if (!target) throw new Error('Usage: node scrape.mjs https://example.com');
const url = new URL(target);
if (!['http:', 'https:'].includes(url.protocol)) {
  throw new Error('Only http and https URLs are allowed');
}

const browser = await puppeteer.launch({headless: true});
try {
  const page = await browser.newPage();
  await page.setViewport({width: 1440, height: 900, deviceScaleFactor: 1});
  await page.goto(url.href, {waitUntil: 'networkidle2', timeout: 30000});
  const result = await page.evaluate(() => ({
    title: document.title,
    text: document.body?.innerText?.slice(0, 20000) ?? ''
  }));
  console.log(JSON.stringify({source_url: url.href, result}));
} finally {
  await browser.close();
}

For production, add redirect and DNS/IP checks, response limits, authentication handling, cancellation and structured error codes. networkidle2 can wait indefinitely on applications that maintain long-lived connections; a bounded timeout and a selector-based readiness check are safer for those sites. Lazy-loaded content may require an explicit scroll operation, and a CAPTCHA or bot challenge should be reported rather than bypassed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP or PDF, while the service handles browser setup for you. See the ScreenshotNeo API documentation for all parameters.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the outcome with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

Other options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size and page ranges, custom CSS or JavaScript, pre-capture clicks, selector or network-idle waits, ad and tracker blocking, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work, easing migration.

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing provides two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to try it without a card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

The client cannot see the scraping tool

Check that the server started successfully, that the transport is configured for the client and that capability negotiation completed. A schema-dialect mismatch or a malformed tool definition can prevent discovery. Inspect startup logs and validate the advertised schema before changing prompts.

Navigation times out

Determine whether the page is waiting for a selector, network idle or an indefinitely open connection. Use a bounded timeout, wait for a page-specific readiness signal and return a partial-result warning when policy allows. Do not solve hangs by granting unrestricted network access.

A redirect is blocked

This is normally the destination policy working as intended. Review the final host and IP, add only the necessary origin to the allowlist and keep private-address blocking enabled. Never accept an arbitrary redirect merely because the first URL was approved.

The result contains instructions for the agent

Classify it as untrusted page content. Preserve the text for analysis, but do not execute its requests, reveal secrets or broaden tool permissions. If the workflow cannot maintain that separation, stop and require a human review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Login state disappears between calls

MCP requests do not create implicit shared state. Use an explicit, expiring session identifier managed by the server, or authenticate each retrieval independently. Keep session cookies isolated per user and origin.

A CAPTCHA or bot check appears

Report the challenge and stop. Do not promise that an MCP wrapper defeats it. For screenshot workflows, ScreenshotNeo marks bot checks and failed loads in its response and does not bill those unsuccessful captures.

FAQ

Can an MCP server make scraping legal?

No. MCP defines communication and tool invocation. Site terms, robots directives, copyright, privacy rules and sector-specific law depend on the target, your purpose and your jurisdiction; obtain permission where required.

Should every page be fetched through a browser?

No. Prefer direct HTTP for stable, public, non-rendered data. Use a browser only when JavaScript rendering, interaction or session state is necessary, because browser execution expands attack surface and resource use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should a scraper return when a page is unavailable?

Return a typed error with the destination, failure class, timeout or HTTP status and whether any partial data is present. Do not convert an error page into authoritative content, and do not silently retry against a different origin.

Frequently Asked Questions

How should pagination be represented in an MCP scraping tool?

Expose an opaque, server-managed cursor with an expiry rather than asking the model to invent offsets or session identifiers. Validate that the cursor belongs to the same user and origin.

Can one MCP server safely serve several teams?

Only with tenant isolation: separate credentials, origin policies, rate limits, logs and session stores, and prevent one tenant’s tool results or browser cookies from entering another tenant’s context.

What is the safest default for a new scraping deployment?

Start with read-only tools, an explicit origin allowlist, no credentials, bounded responses and a restricted container; add browser state or write actions only when a documented requirement justifies them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.