The agents got their own computers
In one quarter, OpenAI, Meta and Anthropic each answered the same question differently: whose machine does the agent run on? That answer predicts your risk far better than desktop-versus-browser does.
Explainers and comparisons on agents that operate a computer. No benchmark scores, for reasons we set out at length.
In one quarter, OpenAI, Meta and Anthropic each answered the same question differently: whose machine does the agent run on? That answer predicts your risk far better than desktop-versus-browser does.
An agent that operates software the way a person does — reading the screen, moving the pointer, typing. Why that matters, what the loop actually looks like, and where it breaks.
The three big labs ship computer use in genuinely different shapes, and at genuinely different stages of maturity — one beta, one GA, one preview. Comparing them on score misses the point.
There is no single best one, and any list that ranks them on a benchmark number is selling you something. Here they are sorted by the job you are actually trying to do.
Ten open-source projects, and one uncomfortable distinction — almost all of them are open harnesses wrapped around a closed model. Only one ships open weights.
Two agents quoted at 38% and 85% on the same benchmark may not have run the same benchmark. A field guide to the variables, and why this directory publishes no scores at all.