Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Computer use is a loop, not a single API call. The model looks at a screenshot, proposes the next click, keypress, scroll, or typed text, and your application decides whether to run it, captures the new screen, and sends that back. The model does not supply the browser session, the desktop, your permissions, or a reliable record of what happened. You build those, and most of the engineering work sits there.
How the loop works
The major providers describe the same basic cycle, though the action format differs. Build these six stages explicitly, because each one is a place where an agent can fail silently.
- Task and policy. Write down the goal, the sites and actions the agent may use, and which actions require a human to approve them before they run.
- Observation. Capture the current screenshot and send it with the task and the relevant conversation state.
- Model request. The model returns its next step. Depending on the integration, that is generated code or a structured action such as click, type, scroll, keypress, wait, or screenshot.
- Execution. Parse and validate the request, enforce access and resource limits, and run it in a controlled browser, desktop, VM, or container.
- Feedback. Capture a new observation and return it to the model.
- Completion check. Stop on completion, refusal, error, or limit. Then check the application’s actual state rather than accepting the model’s account that the job succeeded.
OpenAI documents two patterns. In code execution, the model writes code that your system runs in an isolated environment. In its structured computer tool, the model requests mouse and keyboard input and your application translates those requests into real input events. OpenAI’s guide also names existing UI functions and remote MCP tools as alternatives when your application already exposes higher-level operations. Google’s Computer Use documentation describes a similar client-side loop and uses Playwright as the example browser action handler.
Browser-only or whole desktop?
This is the first architectural decision, because it determines which tool surface you use and how much of the machine you must isolate.
#1 Best Overall
- Browser-only tasks, such as navigating a web app, filling forms, or reading pages, fit a browser runtime. Anthropic documents a separate browser-use tool for this scope and advises using its computer-use tool when a whole desktop is needed.
- Whole-desktop tasks, such as moving between native applications or handling file dialogs, need a desktop or VM the model can see and control. Your harness then owns window focus, file access, and the machine’s lifecycle.
- Tasks with an existing function or API should be checked against direct function calls or MCP tools first. Pixel-level control is the fallback for work that has no better interface.
Provider comparison
These are not interchangeable. The table reflects what each vendor’s published material says at the time of writing, which was early October 2026. Model support, labels, and limits change, so confirm them on the vendor’s current page before you build.
| Item | OpenAI | Anthropic | |
|---|---|---|---|
| Documented surfaces | Code execution in an isolated environment; structured computer tool translated by the application; existing UI functions or remote MCP tools as alternatives | Computer-use tool for whole-desktop tasks; separate browser-use tool for browser-only tasks | Client-side Computer Use loop; Playwright shown as the browser action handler |
| Action format | Generated code, or structured mouse and keyboard requests, depending on the integration | Not specified in this guide; confirm the exact action schema in Anthropic’s tool reference before coding | Not specified in this guide; confirm in Google’s Computer Use docs |
| Availability label | The Operator System Card update dated March 11, 2025 described initial CUA API availability as a limited preview for select developers on tiers 3–5. Treat this as a historical milestone, not current availability. | Compatibility varies by model and platform; check Anthropic’s current compatibility table | Labeled Preview; Google says the capability may contain errors and security vulnerabilities |
| Screenshot guidance | If screenshots are downscaled, the harness must map model coordinates back to the target coordinate space | Model-family pixel limits and a recommended starting size (see below) | Not stated in Google’s Computer Use docs |
| Supervision guidance | Human oversight recommended for OS automation (Operator System Card update, March 11, 2025) | Review and verify actions and logs, and watch for prompt injection through webpages or images | Close supervision for important tasks; advised against critical decisions, sensitive data, or actions where serious errors cannot be corrected |
Runtime state: the conversation is not the browser
The API conversation and the browser or desktop runtime are separate state holders. Continuing the conversation does not restore a browser session, a login, or any runtime variables. Keep the session alive, store tool calls and their results in the conversation, and design explicit behavior for the failures you will meet:
Rank #2
- Timeouts and disconnections: decide in advance whether a run resumes from the last verified step or restarts.
- Stale sessions: detect expired logins and stop for a handoff instead of guessing at credentials.
- Retries: make actions idempotent where you can. A retried “submit” can submit twice.
- Partial completion: record which steps were verified, so a resumed run does not repeat a purchase or a write.
Screenshots and coordinate mapping
Many missed clicks come from a simple mismatch: the model saw a resized image, but the harness clicked in the original coordinate space. OpenAI’s guide warns about exactly this. Anthropic’s best-practices article, dated May 13, 2026, gives model-family limits. Images above a limit may be downscaled internally, so the pixels the model sees and the coordinates it returns must line up.
| Model family (Anthropic, May 13, 2026) | Long-edge limit | Megapixel limit |
|---|---|---|
| Claude 4.6 family | 1568 px | 1.15 MP |
| Claude Opus 4.7 | 2576 px | 3.75 MP |
These are one vendor’s model-specific technical limits. They do not transfer to other providers, and they can change.
Rank #3
- Record the native size. Capture the target display at its native resolution and keep that as the coordinate reference.
- Resize within the limit. Anthropic recommends starting at 1280×720 for most use cases and 1080p for Opus 4.7. A 1920-pixel long edge exceeds the 1568-pixel limit for the 4.6 family, so pre-scale it yourself rather than letting the API do it.
- Store the scale factor with each screenshot.
- Map returned coordinates back. Multiply by native width divided by sent width. For example, a 1920×1080 display sent at 1280×720 has a factor of 1.5, so a model click at (640, 360) becomes (960, 540) on screen.
- Validate bounds before execution, and reject any coordinate outside the display.
Anthropic’s article states: “The single highest impact optimization is also one of the simplest: pre downscale your screenshots before sending them to the API.” This is a vendor recommendation, not an independent benchmark.
Reliability: observe often and verify outcomes
- Return a fresh screenshot whenever the UI state is unknown.
- After a short group of actions, return an observation so the model can confirm the result before it continues.
- Validate action shape and bounds before any coordinate or text reaches the browser or operating system.
- Check the final state in the application itself, not the model’s summary.
Benchmark figures need their context. OpenAI’s Operator System Card update, dated March 11, 2025, reported 38.1% on OSWorld (OpenAI, 2025) for the CUA model in that release context. The same update said the model was not yet highly reliable for OS task automation and recommended human oversight. Read that figure as a dated result for one release. It is not a reliability estimate for your workflow and not a current cross-provider comparison.
Rank #4
Safety controls to build into the harness
Computer-use agents can act on real accounts and data. Put these controls in the harness and the environment, not only in the prompt.
- Isolation: run in an isolated browser, VM, or container, and restrict access to the sites and actions the task needs.
- Untrusted input: treat page, document, and tool-result text as data. Anthropic notes that prompt injection can arrive through webpages or images.
- Confirmations: require a human yes before purchases, data transmission, destructive changes, or typing sensitive information into a form.
- Limits: cap steps, run time, and cost, and provide cancellation and a clear handoff path.
- Audit: log tool activity and review it, including the actions the model proposed and the ones you executed.
- Scope: avoid workflows that need perfect precision or where a mistake cannot be reversed without human supervision.
OpenAI’s guide states the principle directly: “Text in a page, document, or tool result cannot grant permission or override the user’s instructions.”
Quick Recap
Best Value
Troubleshooting common failures
| Symptom | Likely cause | What to check |
|---|---|---|
| Clicks land a few pixels or a large distance from the target | Coordinates were not mapped back after resizing | Compare the stored scale factor with the native and sent dimensions |
| The model reports success, but the record is missing | The model’s final account was trusted instead of verified | Query the application state directly and compare it with the expected result |
| The run resumes but the user is logged out or on a different page | The API conversation was continued without the browser session | Confirm the session is still alive; re-authenticate or restart from the last verified step |
| The same action repeats | No fresh observation after a group of actions, or no step cap | Return a new screenshot and confirm the step limit is enforced |
| Page text instructs the agent to do something outside the task | Untrusted content was treated as an instruction | Do not act on it, block the action, and log the event for review |
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

