Why Synthetic Data Is the Real Bottleneck for Computer Use Agents
Computer use agents, AI that can click, type, scroll, and navigate desktops and browsers, promise a new class of automation. Most teams can get a prototype working quickly, but scaling those agents to production exposes a hard truth: data is the bottleneck. It’s not that models can’t learn the basics, but that real-world interaction data is scarce, expensive, and risky to use. Synthetic data offers a way forward, yet it has a dirty secret: most synthetic datasets are too clean. They miss the noise, the edge cases, and the human-like mistakes that actually matter for agents that need to actually use a computer.
Real data is expensive and risky
Getting high-quality interaction data at scale is hard. To capture realistic workflows, teams often run real agents on production environments and log clicks, keystrokes, screenshots, and system events. The cost is real. A single test of a complex workflow might take hours, require dedicated infrastructure, and risk breaking live systems. Safety is another factor. Recording real user sessions introduces privacy concerns and exposes sensitive data. Even anonymizing data doesn’t eliminate the risk of leaks or regulatory issues. The result: most teams end up with tiny, curated datasets that are barely enough to train a single agent.
Synthetic data is often too clean
Synthetic data promises to solve the scale problem by generating millions of trajectories. The catch is that most synthetic datasets are built with simplified assumptions: perfect clicks, predictable navigation paths, and clean environments. A study on browser automation agents found that synthetic datasets that excluded corner cases led to an average failure rate of 23% on real-world tasks. In contrast, agents trained on a mix of real and synthetic data reduced that failure rate to 8%. The gap shows that synthetic data alone isn’t enough. Agents need to see the messy side of computing: unexpected pop-ups, layout shifts, slow network responses, and the occasional mouse drift.
The gap between simulation and reality
Many teams rely on browser automation libraries or desktop emulators to generate data. These tools are great for controlled experiments, but they don’t capture the full context of a real computer. A simulated browser might not reproduce the exact timing of a network request, the quirks of a specific operating system, or the behavior of third-party extensions. When agents are trained on these synthetic environments, they struggle when they encounter the real thing. The failure modes shift from predictable errors to unpredictable edge cases. This is why synthetic data is the real bottleneck: it looks like data, but it doesn’t behave like real interaction.
The key insight is that synthetic data alone can’t close the gap. You need data that reflects the messiness, variety, and unpredictability of real computer use. The best results come from combining synthetic trajectories with real-world interaction data that covers the edge cases and workflows that matter to your specific use case.
How Coasty fits
Coasty runs computer use agents on real desktops and browsers to capture realistic interaction data. This approach lets teams generate synthetic datasets that reflect the actual complexity of real workflows, including the noise, timing variations, and edge cases that are hard to reproduce in a lab. Coasty’s synthetic data service is custom and contact-led, designed to meet the specific needs of each team. There’s no fixed package or self-serve product, but teams can work directly with the Coasty data team to build datasets that match their use cases, target environments, and success metrics.
The bottleneck isn’t the model, it’s the data. Synthetic data that mirrors real computer use can close the gap between prototypes and production agents. If you’re looking for high-quality interaction data tailored to your workflows, book a data call with the Coasty data team at https://cal.com/coasty/coasty-data-call .