Back to Blog
Tutorial

Michael Rodriguez5 min
Pg Up

Most UI automation tools rely on brittle selectors like IDs, classes, or XPath. When UI changes, tests break. The Coasty computer use API does not guess. It sees the screen and acts. The /v1/ground endpoint is the bridge between visual perception and precise actions. You send a screenshot and a description of what you want, and it returns the exact x,y coordinates to click, type into, or hover over.

How /v1/ground works

Grounding is a two-step process. First, capture a screenshot of the desktop or browser. Second, POST that screenshot with a human-readable description of the target UI element. The request body must include base64-encoded image data, an instruction describing the element, and the cua_version you are using (for v3, send "v3"; for v4, send "v4"). The endpoint returns a JSON response with a status and a coordinate object containing x and y fields. You use those coordinates to perform actions with the same computer use model or another tool.

python
import os
import base64
import requests
import json

def ground_element(image_path, instruction):
    # Read the image file and encode to base64
    with open(image_path, "rb") as f:
        image_bytes = f.read()
    base64_image = base64.b64encode(image_bytes).decode("utf-8")

    # Build the request payload
    payload = {
        "image": base64_image,
        "instruction": instruction,
        "cua_version": "v3"
    }

    # Call /v1/ground (price: $0.03 per request)
    url = "https://coasty.ai/v1/ground"
    headers = {
        "X-API-Key": os.getenv("COASTY_API_KEY")
    }
    resp = requests.post(url, json=payload, headers=headers)
    resp.raise_for_status()
    return resp.json()

# Example: find the "Sign in" button
result = ground_element("screenshot.png", "the sign in button")
print(json.dumps(result, indent=2))

Request and response fields

  • image (string): base64-encoded PNG or JPEG screenshot
  • instruction (string): natural language description of the target element
  • cua_version (string): version of the computer use agent (v3 or v4)
  • status (string): indicates the outcome (e.g., "success")
  • coordinate.x (number): pixel X position relative to the top-left corner
  • coordinate.y (number): pixel Y position relative to the top-left corner

POST https://coasty.ai/v1/ground with image + instruction + cua_version to get x,y coordinates in $0.03.

Where this beats brittle automation

Brute-force selectors rely on a specific HTML structure that can change at any time. These changes break scripts. With /v1/ground, the agent looks at the actual visual representation of the UI. It does not care about the underlying DOM. If the button moves, the screenshot changes, and the ground endpoint recalculates the coordinates automatically. This makes your computer use agent resilient to layout updates, theme changes, and localized UI strings. You build automation that keeps working across releases.

Integrate grounding with a full computer use loop

  • Capture a screenshot of the current state
  • POST that screenshot to /v1/ground to get click coordinates
  • Use those coordinates with click or type actions in your workflow
  • Repeat until the task completes

You now have a reliable way to locate UI elements purely through vision. Use /v1/ground as the eyes of your computer use agent. Build automated workflows that survive UI changes. Get your API key and start grounding at https://coasty.ai/developers.

© 2026 Coasty

Backed byYCombinator