What people actually use them for
Nine concrete tasks, with what the agent does step by step, which tools suit each one, and exactly where each falls over. The demo tasks are here — groceries, flights — but so are the ones that pay.
| Task | Type | Steps | Risk | Reversible |
|---|---|---|---|---|
| Ordering the weekly groceries The demo everyone shows. Fifteen to forty actions, of which two actually require judgement. | Personal | 15–40 | Medium | Partial |
| Booking a flight or a hotel Comparison rather than a fixed sequence, which is what these agents are genuinely good at. | Personal | 20–60 | Medium | Partial |
| Filling in a long application form Pure transcription across several pages. Unglamorous, low-risk, and one of the highest-value things here. | Personal | 20–50 | Low | Yes |
| Working across apps that do not talk to each other A spreadsheet, a PDF in another program, a draft in a third. No API exists between them and none is coming. | Personal | 10–30 | High | Partial |
| Invoice processing The canonical business case. Every organisation has a version of it, and the target system almost never has an API. | Work | 10–25 | Medium | Yes |
| Data entry into a system with no API Where computer use most clearly beats the alternative, because the alternative is a person doing it by hand. | Work | 5–15 per row | Medium | Yes |
| Extracting data trapped in a legacy application Using an agent as a reluctant API for software that never had one. | Work | 10–30 | Low | Yes |
| Testing your own software The safest task here by a distance: your environment, your disposable data, cheap false positives. | Work | 15–40 | Low | Yes |
| Triaging an inbox The highest prompt-injection exposure on this list, and it deserves saying plainly. | Work | 10–40 | High | Partial |
The pattern
Read the nine together and the shape is clear.
The tasks that work are short, verifiable and reversible. Filling a form beats completing a purchase. Extracting data beats writing it. Testing your own software beats operating someone else’s. The demo tasks are shown precisely because they are legible, not because they are the ones that hold up.
The tasks that pay are boring — invoice processing and data entry into systems with no API. Nobody makes a launch video about those, and they are where the money is, because the alternative is a person doing it by hand indefinitely.
And risk tracks the surface, not the difficulty. A browser agent that misreads a fare wastes your time. A desktop agent holding your credentials, reading text written by someone else, is a different category of problem.