Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →To capture a website screenshot with an AI agent in the OpenAI Agents SDK, connect the SDK’s ComputerTool to a browser runtime your application controls. Implement the computer interface—including its screenshot method—then give the tool to an Agent and run it with Runner. The SDK does not provide or host the browser: your application supplies the browser harness.
How the screenshot flow works
ComputerTool bridges the agent to a developer-provided computer or browser implementation. The agent issues computer actions through the tool; your harness performs them in the browser and supplies the current display as a base64-encoded PNG. The relevant boundary is the SDK’s Computer API reference.
- Start a browser runtime in your application and open the requested website in the page or display controlled by that runtime.
- Implement the SDK computer interface for the runtime. Include
screenshot()and the action methods required by the chosen interface, such as clicking, scrolling, typing, waiting, and keyboard input. - Construct
ComputerToolwith that implementation and add it to the agent’s tools. - Run the agent with
Runnerand instructions to navigate to or inspect the requested site. - Have the harness return screenshots in the interface’s required format: base64-encoded PNG data.
The exact setup depends on whether your browser driver is synchronous or asynchronous. The SDK reference defines Computer for synchronous implementations and AsyncComputer for asynchronous ones. Use the form that matches the driver; do not block an async browser workflow by treating it as synchronous.
Use the SDK’s Playwright example as the implementation template
The computer-use guide points to examples/tools/computer_use.py, a Playwright-based harness, as a runnable reference. Follow it for browser startup, implementing the interface, wiring ComputerTool into an agent, and running the agent. The API reference establishes the screenshot return contract, but the code below is intentionally a wiring outline rather than a drop-in Playwright implementation: the interface has multiple required interaction methods, and their exact implementation must match your installed SDK and browser driver.
#1 Best Overall
Do not assume the SDK launches or manages a local Playwright browser for you. The browser process, page lifecycle, and computer methods belong to your application’s harness. Consult the guide and example together, and verify their requirements against the installed Agents SDK version.
Choose sync or async to match your browser driver
- Synchronous driver: implement
Computerand pass that instance toComputerTool. - Asynchronous driver: implement
AsyncComputerand use the corresponding asynchronous methods and execution flow.
In either case, screenshot() must return a base64-encoded PNG of the current display. Returning a file path, JPEG bytes, or raw PNG bytes instead does not meet the documented screenshot contract. See the API reference for method signatures and the guide for the supported workflow.
Check the effective model and computer-use request format
Computer-use request behavior depends on the model that is actually used for the Responses API request. The current guide documents a GA path that sends a computer tool payload and can return batched actions[], and an older computer-use-preview path that uses a computer_use_preview payload and returns one action per call.
Rank #2
Before debugging action handling, check the effective model—not only the model named where the agent is first constructed. A run configuration or prompt template may override it. Model support, defaults, and the GA-versus-preview behavior can change; use the current computer-use guide when selecting a model and confirm that your action loop handles the documented response shape.
Decide whether you need computer use or only a screenshot
ComputerTool is appropriate when the agent must operate a browser through computer-style actions and receive screenshots as the tool cycle proceeds—for example, when it needs to navigate, click, scroll, or inspect the resulting display. If the requirement is only to obtain a screenshot once, this approach still requires you to build and maintain the browser harness and action interface.
The cited SDK documentation describes the ComputerTool path; it does not establish a specific custom-function recipe for returning a one-off screenshot. Avoid treating such a function as an equivalent documented SDK pattern unless you verify its implementation separately.
Rank #3
Troubleshoot common failures
- The tool cannot take a screenshot: Confirm that the harness implements
screenshot()and returns base64-encoded PNG data, as required by the API reference. - Browser actions fail or are unavailable: Check that your implementation supplies the interaction methods required by the selected
ComputerorAsyncComputerinterface. Use the SDK guide’s Playwright example as the implementation reference. - Async calls hang or behave unexpectedly: Match the interface and execution flow to the driver. Use
AsyncComputerfor an asynchronous browser driver rather than mixing synchronous and asynchronous calls. - The model returns a different action shape than expected: Verify the effective model on the actual request and compare it with the guide’s GA and preview paths. A model override can change which format applies.
- The browser never opens the target page: Check the page navigation and lifecycle code in your own harness. The SDK computer tool connects to that implementation; it does not supply the browser runtime.
Or skip the browser setup
If your job is to request a screenshot rather than build an agent-controlled browser, ScreenshotNeo provides a website screenshot API and MCP server. One GET request can return an image or PDF:
See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses include X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for free ScreenshotNeo access.
Frequently Asked Questions
Does OpenAI host the browser used by ComputerTool?
No. Your application supplies and manages the browser runtime and its computer-interface implementation.
What format must the screenshot method return?
A base64-encoded PNG of the current display.
Can the same agent code assume every model uses the same computer action format?
No. The effective model can determine whether the request follows the documented GA or preview format; check the current guide and the model actually used on the request.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

