Search for computer use agent benchmarks and you will find the same handful of models quoted at wildly different scores — figures that differ by more than fifty percentage points, presented with equal confidence, frequently on the same benchmark. They are not all lying. They are mostly not running the same test.
This directory publishes no benchmark scores at all. That is a deliberate editorial position rather than an omission, and this page is the argument for it.
What the benchmarks are
OSWorld is the one most often cited. Built by XLANG Lab, it puts an agent in a real desktop environment — actual applications, actual file system — and scores it on a few hundred open-ended tasks with execution-based verification, meaning it checks the end state of the machine rather than the agent's description of what it did. That design is why it is taken seriously. The original paper is worth reading if you intend to cite numbers from it.
WebArena does something similar for web tasks, in self-hosted clones of real sites so the target does not shift underneath the evaluation.
Both are good benchmarks. Neither produces a number you can lift out of context and put in a comparison table.
Five variables that move the score
1. Step budget
An agent allowed fifty actions per task will complete tasks that an agent allowed fifteen cannot. This single parameter can move a result substantially, and it is frequently unstated in write-ups. A score without its step budget is not a measurement.
2. The harness around the model
Benchmarks nominally evaluate an agent, but an agent is a model plus scaffolding: planning strategy, memory, retry logic, how screenshots are pre-processed, whether an accessibility tree supplements the pixels. Two teams running the same underlying model with different harnesses will report different results, and both may be reported as "model X scores Y".
This cuts to whether the comparison is meaningful at all. If a result reflects a research team's scaffolding as much as the lab's model, then it does not predict what you will get from calling the API yourself.
3. Which variant of the benchmark
OSWorld has been revised, including a verified pass correcting task and evaluation issues found after release. Scores against different variants are not interchangeable, and secondary coverage rarely specifies which was run.
4. The environment
Screen resolution changes what is legible in a screenshot. Operating system version changes application layout. Network conditions change timeouts. These are the sort of details that live in an appendix and are gone by the second retelling.
5. The date
Models are updated continuously, often without a version bump visible to the caller. A figure from six months ago may describe something that no longer exists under that name.
The citation problem
Here is the concrete thing that prompted this policy. Searching for current OSWorld results returns a dense layer of blog posts quoting figures spanning roughly 38% to 86% for overlapping sets of models — several published within weeks of each other, several on the same domain, largely citing one another rather than a leaderboard, and in at least one case attributing a score to a model name we could not corroborate against any primary source.
Meanwhile the actual leaderboard loads its results dynamically, so the pages that scrape it get whatever the scraper saw, frozen and undated.
Reproducing any of those numbers here would be laundering someone else's guess into a table that looks authoritative. A directory's only real asset is being right, and inheriting an unsourced figure spends that asset to fill a column.
What to demand before believing a score
Five things. If a published figure lacks any of them, treat it as an anecdote:
- The number.
- The exact benchmark variant — not "OSWorld" but which revision.
- The step budget.
- The evaluation date.
- A link to the primary leaderboard or paper, not to another article.
Those are also the conditions under which this directory would start publishing scores. The policy is not "benchmarks are useless" — it is that a number without its context is not a benchmark result, it is a rumour with a decimal point.
What to do instead
Benchmark your own task. Ten representative cases from your actual workflow, run against two or three candidates, will tell you more than any public leaderboard. Published benchmarks measure general competence across a task distribution that is almost certainly not yours.
Compare architecture, not accuracy. What it controls, who runs the machine, whether you can embed it — these are stable, checkable and decide most real selections. That is what the directory records, and it is why The best computer use agent for each job is organised by job rather than by rank.
Watch per-step reliability, not per-task. Errors compound: twenty steps at 95% per-step succeeds about a third of the time. When you do run your own evaluation, the useful number is how often a single action is correct, because that is what tells you how long a task can safely get.
Re-run it. These systems are non-deterministic. A single pass over ten tasks has wide error bars, and the difference you think you are seeing between two candidates may not survive a second run.
The uncomfortable part
Refusing to publish scores costs this site traffic. A ranked table is more clickable than an argument about methodology, and the pages quoting unsourced figures will outrank this one for a while.
The bet is that in a field where the vocabulary turns over annually and the numbers turn over faster, being the source that was right is worth more over time than being the source that was first. We would rather have an empty column than a wrong one.
Our full editorial rules are in the repository as a data policy. For the field sorted by task, see The best computer use agent for each job; for the open-source options specifically, Open source computer use agents: what's actually open; and for the mechanics of how these agents work at all, What is a computer use agent?.