Most automation code breaks when buttons change IDs, class names shift, or elements lack selectors. You can wrap a page in a browser and send clicks by XPath, but what if the control is inside a shadow DOM, an iframe, or rendered by JavaScript? You need something that sees the screen and acts like a human. Coasty Computer Use API gives you a vision model that reads screenshots and produces actions. But actions need coordinates. That is where /v1/ground comes in. It turns a natural description like "the download button in the top right" into a precise x,y pair you can feed into pyautogui or your own click logic.
How /v1/ground works
You send a screenshot and a textual description of the UI element you want to locate. The model analyzes the visual context and returns the pixel coordinates of the bounding box. This is a single call, billed at $0.03. The response JSON contains an object with fields x and y representing the center point of the element. It returns only one element per request, so you must be specific about which part you want to ground. The endpoint is designed to work with the other computer use primitives like /v1/predict and /v1/sessions for stateful automation.
curl -X POST https://coasty.ai/v1/ground \
-H 'Content-Type: application/json' \
-H 'X-API-Key: $COASTY_API_KEY' \
-d '{
"screenshot": "$(base64 -i screenshot.png | tr -d '\n')",
"description": "the download button in the top right"
}'
import base64
import os
import requests
api_key = os.getenv('COASTY_API_KEY')
with open('screenshot.png', 'rb') as f:
image_b64 = base64.b64encode(f.read()).decode('utf-8')
resp = requests.post(
'https://coasty.ai/v1/ground',
headers={'X-API-Key': api_key},
json={'screenshot': image_b64, 'description': 'the download button in the top right'}
)
resp.raise_for_status()
result = resp.json()
print(result) # {'x': 1234, 'y': 567}
Key fields and limitations
- screenshot: base64-encoded PNG of the desktop or browser view.
- description: natural language element reference, like 'the search bar' or 'the confirm button in the modal'.
- x and y: pixel coordinates of the element's center, integers.
- Billed at $0.03 per call.
- Returns only the first matching element; if multiple elements match, choose the most specific description.
- Works best with clear, distinct elements; heavily cluttered screens may need refined descriptions.
/v1/ground turns natural language into actionable coordinates for your computer use agent.
Where this beats brittle automation
Traditional automation relies on XPath, CSS selectors, or programmatic IDs. These break when frameworks change, when elements are hidden behind iframes, or when a UI component is rendered dynamically. /v1/ground does not need selectors. It sees the pixel layout and understands human language. You can describe elements by what they look like, where they sit on screen, or how they behave. This makes your agents more resilient to UI churn, better at handling native controls, and capable of working across browsers and desktop apps without brittle selectors.
Ground UI elements to coordinates with /v1/ground and build agents that work on real screens, not just mocked APIs. Combine it with /v1/predict for vision-driven actions or /v1/sessions for stateful automation. Ready to start? Get your API key at https://coasty.ai/developers.
Want to see this in action?
View Case Studies