Guide

Open source computer use agents: what's actually open

Ten open-source projects, and one uncomfortable distinction — almost all of them are open harnesses wrapped around a closed model. Only one ships open weights.

Ten of the twenty agents in this directory are open source. Nine of those ten are open harnesses wrapped around a closed, hosted model. Exactly one ships open weights. If you came here because "open source" meant something specific to you — self-hosting, auditability, no data leaving the building — that distinction is the whole article.

What "open source computer use agent" usually means

Take Browser Use, the most widely adopted project in the category. The library is MIT licensed and you can read every line. Run it, though, and each step of the loop sends a screenshot to Anthropic, OpenAI or Google, and waits for a hosted model to decide what to click.

The harness is open. The intelligence is rented.

That is not a criticism — it is a sensible architecture, and model-agnosticism is genuinely valuable because you can swap providers when pricing or capability changes. But it means the thing people often want from open source here is not what they are getting. You can audit the automation logic. You cannot audit the model, run it offline, or stop screenshots of your work leaving your network.

The one that ships weights

UI-TARS is the exception. ByteDance publishes the model weights under Apache-2.0 alongside a desktop application, so the entire loop can run on hardware you control.

This is the only entry in the directory where "self-hosted" means what an infrastructure engineer would assume it means. For air-gapped environments, regulated data, or anywhere screenshots of the work legally cannot leave the premises, it is effectively the only option in the category — which makes the trade-off worth stating plainly: you need real GPU capacity, and you are exchanging a metered API bill for hardware you buy and operate.

The ten, grouped by what they actually do

Browser harnesses

Browser Use (MIT) hands a browser to any model from Python — the default recommendation for most people starting out. Stagehand (MIT) layers natural-language actions over Playwright in TypeScript, so model calls and deterministic selectors mix in one script. Skyvern (AGPL-3.0) goes vision-first, targeting workflows that break selector-based automation.

Desktop frameworks

Agent S (Apache-2.0) is built around experience — it keeps memory across runs so repeated tasks improve rather than starting cold, which makes it one of the more academically interesting projects here. Open Interpreter (Apache-2.0) predates the current wave and inverts the approach: rather than looking at pixels and moving a pointer, it writes and runs code locally. Self-Operating Computer (MIT) is a minimal reference implementation, valuable mostly for being small enough to read end to end. OpenAdapt (MIT) learns processes from recorded human demonstrations instead of written instructions.

Personal assistants

OpenClaw (MIT) is the odd one out and worth knowing about: a personal assistant from a non-profit foundation that runs on your own devices, reaches you through messaging channels you already use, and can take desktop control. Its documentation is refreshingly blunt about the consequence — tools run on the host for the main session unless you configure sandboxing — which is the correct warning and one most projects bury.

Infrastructure

Cua (MIT) is not an agent at all. It provides containerised macOS and Linux desktops for agents to work inside, solving the sandbox problem that every desktop-scoped deployment eventually hits.

Model weights

UI-TARS (Apache-2.0), as above.

Licences, and the one that will surprise you

Nine of the ten are permissive — six MIT, and Apache-2.0 for UI-TARS, Agent S and Open Interpreter. For most commercial use these impose essentially no constraints.

Skyvern is AGPL-3.0, and that deserves a moment's thought rather than a glance at a badge. The AGPL's network clause extends copyleft to software offered over a network, not just software distributed as binaries. If you build a hosted product on top of an AGPL codebase and let users interact with it over a network, the licence's source-provision obligations are engaged in a way they would not be under MIT or Apache. Skyvern also offers commercial terms, which is the usual resolution.

None of that makes AGPL a bad choice — it is a deliberate one, and it is why open core businesses use it. It does make it a legal question rather than a technical one, and worth raising before the code is load-bearing rather than after.

Every licence above was checked against the project's own repository rather than carried over from secondary sources. We got one wrong in an earlier version of the directory — Open Interpreter was listed as AGPL-3.0 and is in fact Apache-2.0 — which is a reasonable illustration of why the check matters.

Choosing between them

The decision tree is shorter than the field size suggests.

Can screenshots of your work leave your network? If no, the list is UI-TARS — it is the only one where the model itself runs on your hardware. Everything else, OpenClaw included, sends the screen to a hosted model however local the harness is.

Is the task browser-shaped? If yes, take a browser harness — Browser Use for Python, Stagehand if you already run Playwright, Skyvern if selector brittleness across many sites is the specific pain. You inherit the browser's sandbox for free, which is worth more than it sounds.

Does it need the whole desktop? Then the environment question arrives before the agent question, and Cua is the answer to it. Pair it with Agent S or a hosted model's computer use tool.

Is the task genuinely scriptable? Open Interpreter will be faster and more reliable than any pixel-driven approach, because writing code beats simulating clicks whenever code is possible at all.

What open source does not fix

Self-hosting the harness does not change the failure modes. Errors still compound across steps. The screen is still untrusted input, and a locally-run model is just as susceptible to instructions embedded in a web page as a hosted one — the OWASP Top 10 for LLM applications ranks prompt injection first for good reason. Owning the weights changes where your data goes, not what the agent can be talked into.

For nine concrete tasks these tools get pointed at, see the use cases. For the mechanics of the loop and its limits, see What is a computer use agent?. For how the open-source options compare against the hosted labs, see Claude vs OpenAI vs Gemini: three approaches to computer use. The directory filters to open source in one click.