Back to Blog
Research

Sophia Martinez6 min
+W

OSWorld used to be the only game in town for AI computer use benchmarks. Now OSWorld 2.0 has arrived with harder, longer tasks, and the results are embarrassing. AI agents jumped from 12% to 66% task success in a single year, according to the 2026 AI Index Report. But 66% is still garbage when you need an agent that can actually work.

The Benchmark Changed. The Results Didn't Impress.

OSWorld 2.0 was released in June 2026 to address the original benchmark's biggest flaw. OSWORLD 1.0 used binary pass/fail scoring for short tasks, which is too coarse for real workflows. OSWorld 2.0 measures long-horizon tasks that take a human about 1.6 hours. The human baseline on short OSWorld-style tasks sits at 72.4% success. That's the bar you're supposed to clear. Most public model scores hover around 60-70% on OSWorld 2.0, with some top performers barely breaking past human performance. Anthropic's Claude Opus 4.7 with maximum thinking averages a few percentage points higher on OSWorld 2.0. OpenAI's GPT-5.4 reports 75% on OSWorld-Verified. But those are isolated points. The real story is that failure rates compound over time. Claude Opus 4.7 with maximum thinking still fails a non-trivial share of long-horizon tasks. A16Z's analysis of agentic performance found that agents are slower than humans and their failure rates compound over time, making them unreliable for production work.

Why OSWorld 2.0 Is Still Not the Truth

  • OSWorld 2.0 tasks take 1.6 hours on average. That's not a 10-minute task. It's a real workflow.
  • The original OSWorld benchmark saw AI agents jump from 12% to 66% task success in one year. That looks impressive until you compare it to the human baseline.
  • Most vendors only publish cherry-picked highlights. They show 70-75% on OSWorld-Verified or OSWorld 2.0 partial scores while ignoring the rest.
  • OSWorld-Verified is a new verification standard introduced in 2025 by XLANG Lab, but it doesn't fully eliminate contamination or leaderboard gaming.
  • The cost per task varies wildly. One analysis showed a computer-use agent reaching 65.2% partial score at $1.46 per task, while others cost much more for worse results.

The human baseline on short OSWorld-style tasks sits at 72.4% success. That's the bar you're supposed to clear. Most public model scores hover around 60-70% on OSWorld 2.0.

OSWorld 2.0 Proves One Thing: Most Agents Are Not Ready

OSWorld 2.0 was designed to measure realistic long-horizon workflows, and it exposes the weaknesses of current AI agents. Frontier models excel at easy tasks but the success rate plummets as task complexity increases. Agents that can handle a short browser automation task often fail when they have to manage multiple windows, switch contexts, and recover from errors over the course of an hour. The 2026 AI Index Report notes that AI agents made a leap from 12% to ~66% task success on OSWorld, which tests agents on real computer tasks across operating systems, though they still fail. That's the key phrase. They still fail. If you're building production systems that depend on these agents, you're gambling your business on tools that can't reliably complete basic workflows. The vendors know this. That's why they only show partial scores or cherry-pick the easiest tasks for their marketing materials.

How Coasty Actually Delivers on Computer Use

Here's the difference. Coasty.ai is a computer use agent that controls real desktops, browsers, and terminals. Not simulated environments. Not API calls. Real interactions. Our in-house model achieved 85.6% on OSWorld with public results. That's independently verified on the official OSWorld leaderboard at osworld-v1.xlang.ai. We also hit 82.81% on OSWorld. Nobody else is close. Most competitors struggle to break 70%, and even then they're often running on simplified tasks or cherry-picked benchmarks. Coasty runs on desktop apps or cloud VMs and supports agent swarms for parallel execution. You can start for free and bring your own keys (BYOK). We don't hide behind marketing spin. We publish our scores and let you compare them to the human baseline and the rest of the field. If you actually need a computer-using AI that can work, Coasty is the obvious choice.

OSWorld 2.0 shows that AI agents are getting better, but they're still nowhere near ready for production work. Most models hover around the human baseline at best, and failure rates compound over time. If you're betting your business on an AI agent that can't reliably complete a 1.6-hour workflow, you're making a terrible bet. Coasty's 85.6% OSWorld score isn't just a number. It's evidence that real computer use is possible. Try it for free at coasty.ai.

© 2026 Coasty

Backed byYCombinator