Training AI agents that use a computer is hard because real-world interaction data is expensive to collect and risky to share. Companies scan millions of screenshots and logs, but those datasets rarely match the exact workflows and edge cases their models need. Synthetic trajectories can close that gap.
Agents need more than screenshots
Screenshots and logs alone are not enough. An agent needs to understand sequences of clicks, typing, scrolling, and context changes over time. Real clickstreams from legacy systems are messy, incomplete, or unavailable. They also carry privacy risks when they include sensitive credentials or corporate workflows. Synthetic trajectories let teams generate complete, consistent interaction histories that match the target application exactly.
Why trajectories matter for benchmarks
Benchmark suites like OSWorld and WebArena rely on suites of real tasks to grade agent performance. These benchmarks usually contain fewer than 500 trajectories. That small sample size means models are evaluated on a narrow slice of possible workflows. Synthetic trajectories can expand the benchmark set to thousands or tens of thousands of diverse tasks without manual labeling. Research shows that synthetic data can improve generalization when it captures realistic user intent and error patterns.
Techniques that make synthetic data credible
- Computer use agents run on real desktops and browsers to capture realistic input sequences, timing, and error handling.
- Trajectories are filtered to remove noisy or impossible actions, ensuring the model sees only valid workflows.
- Metadata is added to each step, including explanations of user intent, so models can learn causal reasoning, not just patterns.
- Data augmentation techniques such as varying device types, network conditions, and browser states increase diversity without changing the underlying task.
The most valuable synthetic trajectories are those that behave like real users, handling backspaces, retries, and unexpected UI changes, so agents learn to cope with real-world messiness.
How Coasty fits
Coasty runs computer use agents on real desktops and browsers to capture authentic interaction data. This lets the team generate synthetic datasets and trajectories tailored to your specific workflows and constraints. The service is custom and contact-led, meaning you work with the Coasty team to design the data schema, target tasks, and quality criteria that match your goals.
If you need realistic desktop and browser trajectories for agent training or evaluation, book a data call with the Coasty team to discuss your requirements and see how synthetic data can scale your pipeline.
Want to see this in action?
View Case Studies