AWDM: Benchmarking Safe Preparation in Computer-Use Agents under Evolving User Intent
Abstract
Computer-use agents increasingly act on behalf of users in web environments. In real interactions, user intent often takes shape gradually as a task progresses, while agent actions may prematurely commit to choices that have not yet been decided. This raises a central question: how far should an agent proceed before the user has completed the decision process? A well-behaved agent should continue making progress based on confirmed information while preserving choices that remain for the user to decide. We call this capability Safe Preparation. Existing benchmarks mostly focus on one-shot intent clarification or execution under fixed constraints, and provide limited evaluation of whether agents can maintain both task progress and decision boundaries as user intent evolves. To address this gap, we introduce AWDM (Act Without Deciding for Me), a benchmark for evaluating task progress and intent-boundary compliance in multi-stage interactions. AWDM has two core designs. (a) Intent evolution organizes tasks as multi-stage intent sequences in which users progressively add, clarify, or revise their goals, with three controlled levels of intent clarity. This evaluates whether agents can adapt their actions as intent evolves and ask for clarification when critical information is missing. (b) Intent boundaries define the permissible actions at each step based on the intent and permissions disclosed so far and the current environment state, enabling the identification of assumptions or actions that exceed the user's current authorization. Extensive experiments show that final task completion does not reflect stage-wise alignment: agents complete most work items but succeed in only a minority of stages, and they are least safe in tasks that require leaving an open choice to the user. These results point to a persistent gap in preparing without deciding as user intent evolves.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.