Back to Blog
Research

Emily Watson5 min
+Space

OpenAI announced its Computer-Using Agent with a press release about 'breaking new ground.' Then they published the numbers. On OSWorld, the gold standard for AI computer use, Operator scored 38.1%. That is objectively terrible. That is not a computer use agent. That is a broken toy that will delete your data and make you look incompetent to your boss.

The OSWorld numbers nobody wants to talk about

OSWorld 2.0 came out in August 2026. It tests how well an AI actually controls a desktop, clicks, types, and finishes real tasks. The results are brutal. OpenAI's Operator? 38.1%. That's worse than random guessing on many of the tests. Meanwhile, the best models are crossing 80%. We verified our own in-house model at 85.6% on OSWorld with public results, plus 82.81% independently verified on the official leaderboard at osworld-v1.xlang.ai. That gap is not a statistical quirk. It's a reality check.

Microsoft embarrassed OpenAI on browser tasks

  • Microsoft released Fara1.5, a family of browser computer use agents.
  • On the 300-task Online-Mind2Web benchmark, Fara1.5-27B scored 72%.
  • OpenAI Operator? 58.3%.
  • A much smaller open model beat one of the biggest AI companies on the exact thing everyone is hyping up.

Microsoft's Fara1.5-27B scored 72% on Online-Mind2Web, beating OpenAI Operator at 58.3%. This is exactly the kind of story that makes people question whether OpenAI's computer use story is all marketing.

Why your AI productivity gains are disappearing

Workday ran a global survey of 3,200 employees. For every 10 hours of efficiency gained through AI, nearly four hours are lost fixing AI-generated work. That's 40% of your productivity gains vanishing into rework. It gets worse. Gallup's 2026 State of the Global Workplace report says only 20% of employees worldwide are engaged at work. That is a $10 trillion productivity crisis. Companies are rolling out AI computer use tools to fix engagement and productivity, but they're using broken tools that create more work.

The real problem: benchmarks that don't matter

Every week, a new AI model climbs to the top of some benchmark leaderboard. Companies cite these numbers in press releases. Investors use them to justify billion-dollar valuations. But the benchmarks are often incomplete or poorly designed. Anthropic's Computer Use tool gives models screenshots and lets them control a desktop, but the evaluation framework is still evolving. Many vendors test on narrow tasks, browser clicks, simple forms, and pretend that's 'full computer use.' It isn't. Real computer use means handling unexpected errors, broken websites, weird UI layouts, and tasks that require multiple steps over hours, not minutes.

Why Coasty is the only computer use AI that matters

We built Coasty for one reason: we were tired of watching teams deploy broken AI agents that hallucinated, clicked the wrong thing, or failed silently. Coasty is a true computer use agent. It doesn't just call APIs. It controls real desktops, browsers, and terminals. We scored 85.6% on OSWorld with our in-house model using public results, plus 82.81% independently verified on the official OSWorld Verified leaderboard at osworld-v1.xlang.ai. That is higher than every competitor that has published numbers. You can run it as a desktop app, in our cloud VMs, or as agent swarms that work in parallel. It supports BYOK so your data stays yours. There is a free tier if you want to test it yourself. If you're evaluating AI computer use tools and your benchmark scores are below 80%, you're not building automation. You're building a liability.

The era of pretending AI agents work is over. The benchmarks don't lie. OpenAI Operator at 38.1% on OSWorld is a disaster. Microsoft's Fara1.5 beating it on browser tasks is a warning shot. Companies are losing 40% of their AI productivity gains to rework because they're using tools that can't actually do the job. Stop chasing press releases and start looking at real numbers. If you want a computer use agent that actually works, check out Coasty.ai. It's not a toy. It's a tool that will save you time, money, and your reputation.

© 2026 Coasty

Backed byYCombinator