Back to Blog
Enterprise

Rachel Kim7 min
Alt+F4

The finance team needs a monthly reconciliation. A developer builds a bot using UiPath. It clicks a few invoices, sums some values, and sends an email. Six months later, the ERP vendor ships a UI update. The selectors no longer match. The bot fails. A developer has to rebuild the logic. This is the RPA maintenance treadmill. For many organizations, the backlog of such rebuilds outweighs the original automation savings. At the same time, the team has standard operating procedures written in plain English. These documents are already instructions for a human. The gap is that the bot cannot read them directly. It cannot see the screen. It cannot adapt when something unexpected happens. This creates two halves of the same problem. One half is brittle code. The other is unautomated human work.

Why RPA breaks here

Traditional RPA relies on selectors, xpaths, and object IDs. These are brittle references to specific elements on a page. When the application changes, the bot fails. Analysts at Gartner estimate that 60 to 70 percent of a bot’s total cost of ownership comes from maintenance and rework after deployment. This includes updates, bug fixes, and retraining. For many enterprises, the rebuild-on-change cycle is not a one-time event. It is a permanent cost. The bot does not see the screen. It does not know what the user sees. When an exception occurs, the bot halts. A human must intervene. The workflow stops. The process is not auditable beyond what the developer coded. You cannot easily ask, "What exactly did the bot do?" because the bot followed a fixed sequence of actions, not a written procedure. This disconnect makes compliance and debugging hard. You cannot match a human’s steps against a written SOP. You can only match them against the developer’s code.

What changes with computer use agents

  • Agents see the screen like a human. They move the mouse, click, and type. They read the result.
  • They do not depend on brittle selectors or xpaths. When the UI changes, the agent adapts.
  • Explain that they recover from exceptions instead of halting. They can self-correct or escalate when something unexpected happens.
  • They can follow a standard operating procedure written in plain English. The SOP becomes the primary rulebook.
  • They work across any application, including legacy systems, Citrix environments, and virtual desktops where RPA struggles.

A computer use agent follows the SOP you already have, not the code you wrote to approximate it.

How to audit what an AI agent did against the SOP

With a computer use agent, you can build a two-layer audit. The first layer is the SOP itself. You write the steps in clear, plain language. The agent receives this SOP as its primary instruction. The second layer is the agent’s actions on the screen. The agent logs each step it takes: mouse move, click, keystroke, and the text it reads from the screen. You can compare these actions directly to the SOP. For example, the SOP might say, "Open the invoice list, select invoices dated after June 1, 2024, and export them to CSV." The agent logs each of these actions in sequence. If the agent skips a step or makes an error, the discrepancy is visible. You can generate a report that shows exactly which SOP line corresponds to which screen action. This makes it much easier to answer compliance questions. You can show auditors that the process matched the documented procedure at every step. The audit trail is not a black box. It is a readable match between human instructions and machine actions.

How to move without the risk

You do not have to rip out all RPA overnight. The pragmatic path is to start with one high-pain process where the SOP is well-defined and the team struggles with app updates or exceptions. This could be an invoice upload workflow, a customer onboarding checklist, or a monthly compliance report. First, document the SOP in plain language. Second, test the process with a computer use agent in a sandbox environment. Measure how the agent handles the current UI and any common exceptions. Third, compare the agent’s actions to the SOP line by line. Confirm that the workflow is complete and correct. Fourth, integrate the agent into a controlled pilot. Monitor the results, collect feedback, and refine the SOP if needed. Fifth, expand to other processes once you have proven the approach. This phased migration lets you build confidence while keeping legacy bots in place for other use cases. Coasty’s computer use agents run on cloud VMs and offer a desktop app, agent swarms for parallel execution, and a /v1 computer use API. You can start with a free tier to test a few processes without immediate cost commitment.

Why computer use agents are durable

The durability comes from seeing the screen and following the SOP. When the UI changes, the agent adjusts its actions based on what it sees. It does not wait for a developer to update selectors. When an exception occurs, the agent can recover or escalate. It logs the event and continues or stops at the appropriate point. It can also work on legacy systems, Citrix sessions, and virtualized desktops where traditional RPA struggles. This flexibility reduces the long-term cost of ownership. You spend less time rebuilding bots and more time focusing on higher-value improvements. The agent also provides a direct audit trail between SOP and execution. This helps with compliance, debugging, and training. You can hand off a process to a new team member by sharing the SOP and letting the agent execute it, backed by a clear record of actions.

If you are tired of rebuilding bots every time the UI changes, it is time to look at a different approach. Computer use agents see the screen, follow your SOPs, and create an audit trail that RPA cannot match. Book a demo with the Coasty team to see how an agent performs on your own processes and how you can build a durable automation strategy for the long term. https://cal.com/coasty/15min

© 2026 Coasty

Backed byYCombinator