CSS selectors and XPaths are great when they exist, but modern web apps restructure layouts, hide elements, or use dynamic classes. You can end up with a brittle automation that breaks on the next deployment. The /v1/ground endpoint solves this by looking at a screenshot and an element description, then returning the exact pixel coordinates for that element. You can use those x,y values to feed into pyautogui and perform reliable clicks, tabs, or text entry. This gives you a visual grounding layer over any UI.
How /v1/ground works
The /v1/ground endpoint accepts a base64 screenshot and a textual description of the element you want to interact with. It returns a JSON object with the x and y coordinates for that element on the screen. You send the request to POST https://coasty.ai/v1/ground with the COASTY_API_KEY in the Authorization header. The response includes a status field and the coordinates. You can then use these coordinates to perform pyautogui-style actions.
curl -X POST https://coasty.ai/v1/ground \
-H "Authorization: Bearer $COASTY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"screenshot": "$(base64 -i screenshot.png)",
"description": "the submit button with the text submit form"
}'import base64
import os
import requests
api_key = os.getenv("COASTY_API_KEY")
url = "https://coasty.ai/v1/ground"
with open("screenshot.png", "rb") as f:
screenshot_b64 = base64.b64encode(f.read()).decode("utf-8")
payload = {
"screenshot": screenshot_b64,
"description": "the submit button with the text submit form"
}
response = requests.post(
url,
headers={"Authorization": f"Bearer {api_key}"},
json=payload
)
response.raise_for_status()
result = response.json()
print(result)
x = result["x"]
y = result["y"]
print(f"Click at {x}, {y}")Request and response fields
- screenshot (string): base64-encoded PNG image of the screen.
- description (string): natural language description of the target element.
- x (integer): horizontal coordinate of the element's center.
- y (integer): vertical coordinate of the element's center.
- status (string): "ok" for successful grounding, "error" otherwise.
- price: $0.03 per grounding request.
Ground once, click everywhere: after you get x,y, reuse those coordinates for repeated pyautogui operations.
Why this beats brittle selectors
Visual grounding does not rely on CSS classes, IDs, or XPaths. It sees the actual pixel content of the screen, so even if the UI reorders elements or uses dynamic classes, the grounding will still find the element. You do not need to maintain complex selector maps that break with each build. The /v1/ground endpoint works with any screen-based UI, including web browsers, desktop apps, and terminal windows. This makes your automation more resilient to UI changes and reduces the time you spend debugging selector failures.
Add the /v1/ground endpoint to your automation workflow to get reliable coordinates for any UI element. Combine it with pyautogui to perform precise clicks, text entry, and navigation without brittle selectors. Ready to start grounding? Get a key at https://coasty.ai/developers and try the endpoint today.
Want to see this in action?
View Case Studies