You have a desktop app that exposes no API. Maybe it is a legacy Windows tool, a custom Java app, or a web app behind a login form. You need to run a task repeatedly. Traditional RPA tools require fragile selectors. API-only tools only work if the vendor gives you a documented endpoint. The Coasty computer use API solves this by letting an AI agent see the screen and act like a human. It captures screenshots, interprets visual context, and clicks, types, and navigates. You only pay $0.05 per agent step. No brittle selectors, no manual effort.
How it works
The Coasty computer use API runs a task by repeatedly capturing a screenshot, predicting an action, and executing it. You start a task run with POST /v1/runs. The server provisions a cloud machine, launches the desktop, and drives it. Each prediction step costs $0.05. The response includes the status (queued, running, awaiting_human, succeeded, failed, cancelled, timed_out). You can stream events with GET /v1/runs/{id}/events to see progress in real time. If the agent hits a human interaction point, you can pause, resume, or cancel with /v1/runs/{id}/cancel or /v1/runs/{id}/resume. For more control, you can also use the stateful trajectory API: POST /v1/sessions creates a session, then POST /v1/sessions/{id}/predict uses visual context and memory to generate actions at $0.04 per call.
curl -X POST https://coasty.ai/v1/runs \
-H "X-API-Key: $COASTY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"machine_id": "my-prod-desktop",
"task": "Open Chrome, navigate to https://coasty.ai, click the Developers link in the footer, and take a screenshot of the page."
}'Key fields and options
- machine_id: identifier of the cloud desktop (provisioned with POST /v1/machines)
- task: natural language instruction the agent follows
- cua_version: either 'v3' (guided) or 'v4' (autonomous with pass/fail verifier)
- instructions: optional text to append to the base prompt
- system_prompt: custom system instructions for the agent
- max_steps: maximum prediction steps (default and max depend on plan)
- deadline_seconds: task timeout in seconds
- on_awaiting_human: 'pause', 'fail', or 'cancel' when the agent needs human input
- webhook_url: optional URL to receive run state updates (HMAC signed)
- budget_cents: optional budget cap in cents
- States: queued, running, awaiting_human, succeeded, failed, cancelled, timed_out
Start any task with POST /v1/runs and stream events with GET /v1/runs/{id}/events to see the agent in action.
Where this beats brittle automation
Traditional automation relies on selectors like XPath, CSS, or ID. If the UI changes, the script breaks. The Coasty computer use agent sees the screen like a human. It understands context and adapts to layout changes. It can click buttons by text, fill forms by field labels, and navigate by visual elements. You do not need to maintain fragile selectors. The agent can also handle dynamic content, popups, and multi-step workflows without extra code. You only pay per successful step, not per line of code. This makes it practical to automate complex, undocumented desktop tools that would otherwise be impossible.
Next steps
- Provision a machine with POST /v1/machines to get a cloud desktop identity.
- Create a task run with POST /v1/runs and watch progress via GET /v1/runs/{id}/events.
- Use the free POST /v1/parse endpoint to convert existing pyautogui scripts into structured actions.
- Integrate the stateful trajectory API with POST /v1/sessions and POST /v1/sessions/{id}/predict for long-running workflows.
- Drive Coasty from Cursor, Claude Desktop, or other apps via the MCP server.
- Get a key and start automating at https://coasty.ai/developers.
You can now automate any desktop app that has a graphical interface. No more brittle selectors. No more manual scripts. Start a task run with POST /v1/runs, stream the events, and watch the agent complete the work. Build desktop agents that see, think, and act. Get your API key at https://coasty.ai/developers.
Want to see this in action?
View Case Studies