Comparison

Claude vs OpenAI vs Gemini: three approaches to computer use

The three big labs ship computer use in genuinely different shapes, and at genuinely different stages of maturity — one beta, one GA, one preview. Comparing them on score misses the point.

The three large labs all ship computer use, and the most common way of comparing them — which one scores higher — is close to meaningless, because they are not the same kind of product. One is a beta tool inside an API, one is a finished consumer feature, one is a preview model on a cloud platform. Choosing between them is mostly a question about architecture and maturity, not capability.

The shapes are different

Before any capability question, understand what you are actually buying.

Claude Computer Use is a tool inside the Anthropic Messages API. There is no product. The model receives screenshots and returns actions, and you write the loop that executes them against a desktop you supply and secure. Anthropic documents the tool; the machine is your problem. It remains a beta feature and requires a beta header on the request, which is worth knowing before you build a roadmap on it. (Anthropic documentation)

OpenAI ships two things, which is a frequent source of confusion. ChatGPT Agent is the consumer feature inside ChatGPT: a hosted virtual browser with an attached terminal, aimed at people who want a task done. Computer use tool is the same capability offered to developers, where the control loop is yours again — and it is now generally available. The older computer-use-preview path still referenced in a lot of tutorials is deprecated; OpenAI's guidance is to keep it only for existing integrations. (OpenAI platform docs) The earlier standalone Operator has largely been folded into the ChatGPT product.

Gemini Computer Use is a computer use model on the Gemini API and Vertex AI. It launched browser-first, but Google now documents it across browser, mobile and desktop. It is a preview capability, and Google's own documentation warns it may contain errors and security vulnerabilities. (Gemini API docs) Separately, Project Mariner is a DeepMind research prototype that browses on the user's behalf, gated behind the AI Ultra tier in the US, with its capabilities migrating into the Gemini API.

The axis that actually decides it

Not accuracy. Three questions:

1. What does it need to touch?

This axis has narrowed. All three now document desktop as well as browser control, so the question is less "which one can" and more "which one should".

If your task lives entirely in a browser, constrain the agent to a browser anyway — not as a compromise, but because the browser is already a sandbox with well-understood failure modes and cheap state reset. Choosing a desktop-scoped deployment for a browser-only task means accepting a much larger blast radius for no benefit, and that logic holds whichever of the three you pick.

If your task needs a native desktop application — the accounting package, the terminal, a file dialog — Claude was built for that surface from the start, and both OpenAI and Google now document it too.

2. How much churn can you absorb?

Worth pricing in, because the three are at genuinely different stages. OpenAI's developer tool is generally available. Anthropic's is a beta requiring an explicit header. Google's is a preview its own documentation cautions about. That ordering is a snapshot rather than a verdict — it will change — but it should weigh on anything you intend to keep running.

3. Who runs the machine?

This is the question teams underestimate, and it is where most of the real cost sits.

With Claude, and with OpenAI's API model, you provide the environment. That means a virtual machine, a snapshot-and-reset story, network egress rules, credential scoping, and a plan for what happens when the agent does something unintended. That is genuine infrastructure work, and it is the reason Cua exists as a standalone project.

With ChatGPT Agent, OpenAI runs it. You give up control of the environment and gain the fact that securing it is not your job. For a lot of use cases that is the right trade, and pretending otherwise is engineering vanity.

A comparison table that is honest about what it knows

 Claude Computer UseChatGPT AgentOpenAI API modelGemini Computer Use
ShapeTool in an APIConsumer productTool in an APIModel in an API
StageBetaGAGAPreview
ControlsFull desktopHosted browser + terminalDesktop and browserBrowser, mobile, desktop
EnvironmentYoursOpenAI'sYoursYours
Built forDevelopersEnd usersDevelopersDevelopers
EmbeddableYesNoYesYes

You will notice there is no accuracy row. That is deliberate, and it is the most important thing on this page — see Why computer use benchmark scores don't compare for why a single published number comparing these three would be misleading rather than helpful.

What about the three-way comparisons people search for?

A recurring search in this niche asks for Claude versus OpenAI versus a third lab — often xAI. We have left that out rather than fill the column in. The directory's rule is that every claim links to primary vendor documentation, and where a documented, generally available computer use offering of the same kind does not exist, the honest answer is to say so instead of inventing a row. If that changes it will appear in the directory first.

How to actually choose

Pick ChatGPT Agent if the user is a person with a task, not a system with a workload. It is the only one of the four that a non-developer can use directly, and the hosted environment removes the hardest part of the problem.

Pick Gemini Computer Use if you are already on Google Cloud and want one model across web and native interfaces. Weigh the preview status: Google's own documentation warns it may contain errors and security vulnerabilities, which is a reasonable thing to ship a prototype on and a poor thing to ship a product on.

Pick Claude Computer Use if you need the whole desktop and have the appetite to run the environment properly. The full surface is the differentiator and the liability in the same breath. Note the beta header requirement.

Pick OpenAI's computer use tool if you are building your own harness and want the control loop. Of the three it is the only one at general availability on the developer surface, which matters if you are committing a roadmap to it.

Consider none of them if the task is browser-shaped and you want to own the stack. Browser Use and Stagehand are model-agnostic harnesses — you can run any of these labs' models underneath, and swap when the economics change. See Open source computer use agents: what's actually open.

For the wider field sorted by task rather than by vendor, see The best computer use agent for each job. If you are new to the category, What is a computer use agent? covers how the loop works and where it breaks.