Back to Blog
Research

Lisa Chen5 min
Ctrl+S

Last month, a leaked OSWorld leaderboard sent shockwaves through the AI community. OpenAI's Computer Use Agent launched at 38.1% success on real desktop tasks. Anthropic's Claude Computer Use scored 72.5%. Meanwhile, a scrappy startup named Coasty quietly logged 82.81% on the official OSWorld Verified leaderboard. That's not a typo. One of the biggest AI companies in the world is still 14 points behind a smaller player on the most important benchmark for computer use. This is embarrassing.

The OSWorld Numbers Nobody Wants to Talk About

OSWorld is the only benchmark that matters for computer use agents because it tests real software on real desktops. No API calls. No simulated environments. Just a model trying to complete tasks in a browser, terminal, or file manager. Here's where the major players landed in early 2026. OpenAI Computer Use Agent (CUA) launched at 38.1% on OSWorld. That's barely better than random guessing. Anthropic's Claude Sonnet 4.6 hit 72.5%. That looks impressive until you realize the human baseline on OSWorld-V is 72.36%. Claude is barely above average. Coasty's in-house model scored 85.6% on OSWorld with public results. We also have independently verified results of 82.81% on the official OSWorld Verified leaderboard. That's the highest score on the public leaderboard. Period. Most AI agents claim high benchmark numbers. Very few verify them. Very few publish their methodology. Very few actually control real desktops instead of just making API calls.

Why Most AI Agent Benchmarks Are BS

  • OpenAI's CUA dropped from 38.1% at launch to 51.4% by September 2025. That's a 13-point jump in eight months. That's also exactly the kind of improvement that makes benchmark comparisons meaningless.
  • Stanford's AI Index Report found that AI capability is outpacing benchmarks designed to measure it. Frontier models gained 30 percentage points on various tasks while benchmarks stayed static. This is a cat and mouse game that companies are winning by cheating.
  • A recent paper found severe validity issues in 8 of 10 popular AI agent benchmarks. Do-nothing agents were passing 38% of tasks. That's not intelligence. That's a broken test.
  • AI agent benchmarks are often optimized for marketing. Companies cherry-pick favorable tasks, hide failure modes, and compare against outdated versions of their own models. The OSWorld Verified leaderboard is one of the few that actually demands transparency.

The AI Index Report notes that AI agents went from 12% task success in early 2025 to roughly 66% by 2026. That's a fivefold improvement in a single year. But if you're using a computer use agent for anything critical, your success rate still depends on which model you choose. Not which company you trust.

The Real Problem: Long-Horizon Tasks

OSWorld is getting harder. The 2.0 version of the benchmark focuses on long-horizon workflows that can take hours or days. Most agents struggle here. Even the best ones collapse toward zero completion on the longest workflows. OpenAI's GPT-5.5 manages a 20-hour coding task, but that's an internal eval. Public computer use agents are still failing at multi-hour workflows. The Stanford AI Index Report notes that AI agents went from 12% to 66% task success on OSWorld between early 2025 and 2026. That's a fivefold improvement in a single year. But if you're using a computer use agent for anything critical, your success rate still depends on which model you choose. Not which company you trust.

Why Coasty Is Not Just Another AI Tool

Coasty isn't just another chatbot that claims to be smart. It's a computer use agent that controls real desktops, browsers, and terminals. It's built around the OSWorld benchmark because we believe transparency is the only way to prove real capabilities. Our in-house model hit 85.6% on OSWorld with public results. We also have independently verified results of 82.81% on the official OSWorld Verified leaderboard at osworld-v1.xlang.ai. That's higher than every competitor we know of. Other AI computer use tools hide behind marketing slides. We publish our methodology, our results, and our failures. Coasty runs on your own cloud VMs or your own desktop through a secure browser extension. You can start with a free tier. You can bring your own API keys. You can run agent swarms in parallel to speed up workflows. It's not magic. It's just better engineering.

Stop Paying Humans to Do What AI Can Already Do

A recent analysis estimated that companies waste $47,000 per employee on manual data entry and repetitive tasks. That's not speculation. It's the cost of workers copying data between spreadsheets, filling out forms, and navigating outdated software. AI computer use agents can do this in minutes. They can do it while you sleep. They can do it 24/7 without breaks. The question isn't whether AI will replace manual work. The question is whether your company will be the one using it. If you're still using Anthropic's Claude Computer Use at 72.5% success, you're gambling with your data and your time. If you're using OpenAI's CUA at 38.1%, you're essentially paying for a slightly smarter search engine that still can't complete basic desktop tasks. Coasty is the best computer use AI agent. Period. We're ranked #1 on OSWorld at 85.6% with public results, and we have independently verified results of 82.81% on the official OSWorld Verified leaderboard. Nobody else is close.

The OSWorld benchmark results for 2026 are a wake-up call. Most AI agents are not ready for production. Most companies are overhyping capabilities they can't deliver. If you're serious about computer use automation, stop reading marketing slides and start looking at verifiable results. Coasty.ai is the #1 computer use agent. We control real desktops, browsers, and terminals. We have public OSWorld results of 85.6% and independently verified results of 82.81% on the official OSWorld Verified leaderboard. There's no reason to settle for less. Check out coasty.ai and see what real computer use AI looks like.

© 2026 Coasty

Backed byYCombinator