acceptodds
Under review as a conference paper at ICLR 2027

Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents

Abstract

Computer-use agents (CUAs) perform tasks across applications (such as desktops, mobile apps, and web browsers) by observing graphical interfaces and issuing commands such as clicks and keystrokes. These interfaces combine trusted controls and content with untrusted content needed for legitimate tasks. An adversary controlling this untrusted content can embed instructions or misleading visual cues to change the agent's intended action or redirect its commands to the wrong interface target. We formalize security requirements for both the agent's decisions and their execution through GUI commands. Using an ideal execution model, we show that enforcing these requirements at each step protects complete execution traces. We instantiate this model in , our system for secure CUA execution. Its key idea is to commit to an explicit per-action program, called an *action transaction*, before accessing untrusted content. Each transaction fixes its queries to untrusted content and the permitted uses of their responses. The system masks untrusted regions and evaluates each transaction to produce the next action, using an isolated query model to answer its queries. It then locates the intended interface target using the masked interface. Under the model's assumptions, is secure by design, while generating a new transaction at each step enables adaptation to changing interfaces. We evaluate under benign conditions on 400 WebArena tasks using three frontier models across seeds, yielding 6,000 execution traces. achieves an average task success rate of %, compared with % for Vanilla-CUA and % for CaMeL-CUA.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.