A computer use agent is an AI system that operates software through the same interface a person does — it looks at the screen, decides what to do, and moves the pointer or types. It does not call an API. It uses the application.
That distinction sounds pedantic and is in fact the entire point. Almost every other way of automating software requires the software to cooperate: an API, a webhook, a database you are allowed to write to, an export format. A computer use agent requires none of that. If a human can operate it, the agent can attempt it.
This is why the category exists despite being slower, costlier and less reliable than an API call. Most software in the world has no useful API. The internal expenses tool, the vendor portal your supplier insists on, the twenty-year-old desktop application that runs a factory — none of them are going to grow one.
The loop
Strip away the marketing and every computer use agent runs roughly the same cycle:
- Capture the current state of the screen, usually as a screenshot.
- Send it to a multimodal model along with the goal and the history so far.
- The model returns an action — click at these coordinates, type this string, scroll, wait.
- Execute the action against the real machine.
- Capture the screen again. Repeat until done or stuck.
Everything interesting about a given agent is a variation on that loop. Does it see raw pixels or an accessibility tree? Does it plan several steps ahead or react one frame at a time? Does it remember what happened on the last run? Who is responsible for the machine being driven?
Claude Computer Use exposes this loop directly: the model returns actions and you write the code that performs them, against a desktop you provide. ChatGPT Agent hides it entirely behind a product. Both are computer use agents; they sit at opposite ends of how much of the loop you own.
What people call it
The vocabulary has not settled, which is worth knowing if you are searching for prior art.
- Computer use is Anthropic's term, and the one that has spread furthest as a generic label.
- CUA — Computer-Using Agent — is OpenAI's, and also the name of the model behind its agent products.
- GUI agent is what most of the academic literature says.
- Browser agent is the narrower subcategory that only drives a web browser.
They are not quite synonyms, and the last one matters most in practice.
Browser or desktop
The single most useful question to ask about any agent in this category is what it is allowed to touch.
Browser-scoped agents drive a web browser and nothing else. That is a real limitation, but it buys an enormous amount: the browser is already a sandbox, the failure modes are well understood, and tab state is easy to reset. The most widely used open-source options are scoped this way — Browser Use and Stagehand both drive a browser and stop there.
Desktop-scoped agents drive the whole operating system. They can reach the accounting package and the terminal and the file manager — and so can a prompt injection that talks them into it. Claude Computer Use and Agent S work at this level. So does UI-TARS, which is unusual in shipping open model weights rather than calling a hosted service.
The line between the two is blurrier than it was. Gemini Computer Use launched browser-first but is now documented across browser, mobile and desktop, and Computer use tool ships harnesses for both a browser and a full virtual machine. Check the current documentation rather than assuming a product's original scope still holds.
If your task lives entirely in a browser, using a desktop agent buys you risk you do not need.
There is a third shape worth knowing about, which is to stop bolting an agent onto a browser and make the browser itself agentic. Comet takes that route: the assistant lives in the browsing surface, with the page you are on as its working context. It is a product rather than something you can build on, but architecturally it is a distinct answer to the same question.
Where it breaks
Four failure modes account for most disappointment with this technology, and none of them are going away soon.
Errors compound. A task needing twenty steps at 95% per-step reliability succeeds about a third of the time. Per-step accuracy has to get very high before long tasks become dependable, which is why the honest deployments today are short and well-bounded rather than sprawling.
Every step costs a model call. A screenshot is a large input, and you send one per step. Cost and latency scale with the number of actions, not the value of the task, so a slow twelve-step form fill costs the same whether it saved you five minutes or five seconds.
The screen is untrusted input. This is the one that should worry you. An agent reading a web page cannot easily distinguish your instructions from text on the page instructing it to do something else. Prompt injection through rendered content is a live problem — the OWASP Top 10 for LLM applications lists it first — and a desktop-scoped agent with your credentials is a considerably more attractive target than a chatbot.
Interfaces move. Agents that work from pixels are more robust to a changed CSS class than a selector-based script, but a redesigned page still breaks them, and it breaks them at run time rather than at test time.
How it differs from RPA
Robotic process automation has driven GUIs for two decades. The difference is brittleness versus judgement. Classic RPA replays a recorded, deterministic script: fast, cheap, reliable, and broken the instant a button moves. A computer use agent re-derives what to do from what it sees, so it survives a layout change and can handle a case nobody scripted — at the cost of being slower, more expensive, and non-deterministic.
Neither replaces the other. A stable high-volume process still wants RPA or an API. The agent earns its place on the long tail of tasks that were never worth automating because scripting them cost more than doing them.
OpenAdapt sits interestingly between the two — it learns a process from a recorded human demonstration rather than from written instructions, which sidesteps the hardest part of the problem: describing a GUI workflow precisely enough in prose.
Where to start
If you want to try one this week, the shortest path is a browser-scoped open-source harness — Browser Use if you write Python, Stagehand if you already have Playwright tests. Both let you point a model at a browser in an afternoon, and both make the limits obvious quickly, which is the fastest way to learn whether this technology fits your problem.
If you need full desktop control, the question becomes who runs the machine, and that is an infrastructure decision before it is a model decision. Cua exists specifically to solve it.
For what these agents get pointed at in practice — ordering groceries, booking travel, processing invoices, and where each of those falls over — see the use cases.
The full directory lists every agent we track with its control surface, licence and access model. For choosing between them by task, see The best computer use agent for each job. For why you will not find benchmark scores anywhere on this site, see Why computer use benchmark scores don't compare.