Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use WebDriver BiDi’s browsingContext.captureScreenshot command. With an active BiDi session and a browsing-context ID, send the command to receive a Base64-encoded PNG by default. Omit origin for the visible viewport; set origin to document to capture the entire scrollable page. This is an automation protocol command documented by MDN, not a standalone browser JavaScript function named “MDN Screenshot API.”

What the MDN screenshot API actually is

MDN’s browser-automation screenshot method is the WebDriver BiDi command browsingContext.captureScreenshot. WebDriver BiDi is a bidirectional protocol: your automation client maintains a connection to the browser, creates a session, opens or selects a browsing context, and sends protocol messages. The command cannot be pasted into an ordinary page console without that connection and an active session.

The command returns image data encoded as Base64. Your client must decode that value and write the resulting bytes to a file or display them. PNG is the default format. You can request another supported image MIME type, such as JPEG, and provide a lossy-quality value from 0.0 to 1.0.

Viewport, full page, and clipped captures

Visible viewport

Send the command with the active context ID and no origin property. The browser captures what is currently visible in that context’s viewport. Content below the fold is not included.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Full scrollable document

Set origin to document. The result includes the scrollable document, including content outside the current viewport. This is the option to use for a long article, dashboard, or landing page when you need one image rather than a viewport-sized slice.

One element

Pass a clip whose type is element and whose value is the element’s shared ID. Obtain that ID with browsingContext.locateNodes, script.evaluate, or script.callFunction. The element must belong to the document in the context you are capturing.

Rectangular crop

A rectangular clip lets you specify offsets and dimensions instead of an element. This is useful when the target is defined by coordinates or when you need a fixed crop. MDN’s example uses an element’s bounding box, including when the element has been scrolled out of view.

Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Capture workflow with WebDriver BiDi

  1. Start a BiDi-capable browser session. Your automation library must expose a WebDriver BiDi connection and an active session.
  2. Identify the browsing context. Use the context ID returned by your session or navigation commands. A context ID is not the page URL and cannot be guessed from it.
  3. Navigate and wait for the page state you need. If the page is still changing, the screenshot can contain incomplete content. Wait in your client for the relevant navigation or application condition.
  4. Send the command. The protocol method is browsingContext.captureScreenshot; include the context and any origin, format, quality, or clip options.
  5. Decode the response. Read the returned Base64 image string, decode it, and save it with an extension matching the selected format.

Viewport PNG message

{
  "id": 7,
  "method": "browsingContext.captureScreenshot",
  "params": {
    "context": "YOUR_CONTEXT_ID"
  }
}

Replace YOUR_CONTEXT_ID with the ID from your active session. The response contains the encoded image data; the exact envelope is handled by your BiDi client.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Full-page JPEG message

{
  "id": 8,
  "method": "browsingContext.captureScreenshot",
  "params": {
    "context": "YOUR_CONTEXT_ID",
    "origin": "document",
    "format": {
      "type": "image/jpeg",
      "quality": 0.8
    }
  }
}

JPEG quality is meaningful for lossy formats. If you omit quality, the browser chooses the compression behavior. Use PNG when pixel-accurate text, transparency, or lossless output matters; use JPEG when a smaller photographic image is more important.

Element clip shape

{
  "id": 9,
  "method": "browsingContext.captureScreenshot",
  "params": {
    "context": "YOUR_CONTEXT_ID",
    "clip": {
      "type": "element",
      "element": {
        "sharedId": "ELEMENT_SHARED_ID"
      }
    }
  }
}

The shared ID must be obtained from the same document and context. For a coordinate crop, send the protocol’s rectangle clip with the required offsets and dimensions instead of an element clip.

Decoding and saving the Base64 result

Your automation library normally exposes the result as a string. Decode that string as Base64 before writing it. In a language with a binary file API, open the output in binary mode; writing the encoded text directly creates a corrupt image.

// Conceptual JavaScript handling after your BiDi client returns the result
const bytes = Buffer.from(result.data, 'base64');
require('fs').writeFileSync('page.png', bytes);

The property name used for the encoded data can differ between BiDi libraries, so consult the library’s response object. The protocol value itself is Base64 image data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choosing WebDriver BiDi or the Screen Capture API

Need Use What happens
Automated screenshot of a context or document browsingContext.captureScreenshot Your automation client sends a BiDi command and receives encoded image data.
A person chooses a tab, window, or monitor to share getDisplayMedia() The browser presents a selection UI and returns a live media stream.
Still image from a selected display stream ImageCapture.grabFrame() plus canvas encoding Capture begins as a user-approved stream, then a frame is drawn and encoded.

getDisplayMedia() is not a silent website screenshot function. It requires the browser’s user-selection and permission flow, and recent user interaction (transient activation). Permissions Policy can gate display capture through the display-capture directive in an HTTP header or an iframe’s allow attribute, but policy permission does not remove the user prompt.

Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

Element Capture versus Region Capture

These are controls for display streams, not replacements for the BiDi screenshot command. Element Capture restricts the stream to a selected DOM tree and its descendants, excluding content outside that tree. Region Capture uses a DOM tree’s bounding box in the tab; overlapping content can appear over the intended region. Choose Element Capture when private material outside the target tree must be excluded, and Region Capture when the tab area itself is what you want.

getDisplayMedia() has limited availability and is not Baseline according to MDN. Check the current compatibility table before promising support for a particular browser. Do not infer the same support or security behavior for WebDriver BiDi’s screenshot command without checking the implementation you deploy.

Common errors and fixes

Error Likely cause Fix
invalid argument A required parameter is missing or has the wrong type. Verify that context is a string ID and that origin, format, and clip follow the protocol shapes.
no such element The shared element ID cannot be resolved or belongs to another document/context. Locate the node again after navigation and use the ID returned for the captured context.
no such frame The context ID is unknown. Use the context created by the current session; do not reuse an ID after the session ends.
unable to capture screen The requested clip intersects the origin with zero width or height. Check the rectangle dimensions and element bounds; ensure the target has non-zero rendered size.
unsupported operation The browser cannot capture that context. Confirm that the selected browser and context support the operation, then try a supported context or browser version.

Blank or incomplete output

  • Wait for navigation and application rendering before issuing the command.
  • For a long page, use origin: "document" rather than assuming a viewport capture is full page.
  • Confirm that your decoder writes binary bytes, not the Base64 characters.
  • When clipping an element, re-locate it after any navigation or DOM replacement.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, format, and reliability considerations

Full-document captures can be substantially larger and slower than viewport images because more pixels must be rendered and encoded. Capture only the origin and area you need. Use JPEG quality when a smaller lossy file is acceptable; retain PNG for sharp interfaces, transparency, and archival comparisons.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make your automation wait on a deterministic condition rather than an arbitrary short delay. Record the URL, context ID, selected origin, format, and clip in your job logs so a failed image can be reproduced. Treat a screenshot as a point-in-time rendering: animations, delayed content, personalization, and network failures can change the result between runs.

Or skip the browser setup

ScreenshotNeo provides a single-request website screenshot API when you do not want to maintain a BiDi session. It removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server gives AI agents tools named take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for parameters and response details. A cURL request:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Responses identify the page verdict and billing status with X-Page-Verdict and X-Billed headers. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account to begin.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Practical decision checklist

  • Choose BiDi when your test or automation already controls a browser session and needs a protocol-level capture.
  • Choose origin: "document" for a full scrollable page; leave it out for the viewport.
  • Use an element shared ID for semantic element capture and a rectangle for coordinate-based crops.
  • Use getDisplayMedia() only when a person must select a display surface and approve sharing.
  • Use ScreenshotNeo when a hosted request, cleaned capture, billing verdicts, or AI-agent integration is more useful than browser-session management.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.