Back to Blog
Tutorial

Sarah Chen6 min
End

Most automation tools rely on brittle selectors and abstractions. They break when layouts change or when you need to click something that does not map cleanly to an API. The Coasty /v1/predict endpoint flips that by letting you send a raw screenshot and a natural language instruction, then receiving a list of actions that a human could execute. You loop capture, predict, act until the status is done. This gives you a true computer use agent that sees and reasons like a person.

How it works

The endpoint is POST https://coasty.ai/v1/predict. It expects a JSON body with three fields. Screenshot is a base64-encoded PNG image. Instruction is a string describing what to do. cua_version is a string like v3 that tells the model which computer use version to use. The request is charged $0.05. The response is a JSON object with actions and status. Actions is an array of objects, each with type (e.g., click, type), x, y, and optional text. Status is one of running, done, or an error code. You call predict repeatedly until status is done.

bash
#!/bin/bash
# Example using curl. Replace <YOUR_KEY> with your actual key.

KEY=$(cat "$HOME/.coasty_key")

# Base64 encode the screenshot (replace with your file)
SCREENSHOT=$(base64 -i screenshot.png)

# Call /v1/predict
RESPONSE=$(curl -s -X POST https://coasty.ai/v1/predict \
  -H "X-API-Key: $KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"screenshot\": \"$SCREENSHOT\",
    \"instruction\": \"Click the 'Submit' button\",
    \"cua_version\": \"v3\"
  }")

echo "Response: $RESPONSE"

Stateless loop with capture, predict, act

  • Send a screenshot and instruction to /v1/predict.
  • Receive actions and a status flag.
  • Apply actions (e.g., pyautogui.click, pyautogui.write).
  • Capture a new screenshot only when status is running.
  • Stop when status is done or an error occurs.
  • Each predict call costs $0.05.
  • No need to maintain a session or session memory here.

Loop capture → predict → act until the status is done.

Where this beats brittle automation

Traditional automation might look for an element by ID or class. If the UI changes slightly, the selector breaks. With the /v1/predict endpoint you are letting the model see the current state of the screen and reason about it. It can handle new buttons, different layouts, or ambiguous elements because it works directly with pixel data and natural language. This is the core of a computer use agent: real visual perception instead of brittle selectors.

You now know how to turn screenshots into actions with the /v1/predict endpoint. Build an agent that sees, reasons, and clicks just like a human. To get a key and start building, visit https://coasty.ai/developers.

© 2026 Coasty

Backed byYCombinator