acceptodds
Under review as a conference paper at ICLR 2027

AlphaCUA: Scaling Safety Benchmarks for Computer-Use Agents by Discovering and Validating Risk Factors

Abstract

Computer-use agents (CUAs) execute high-stakes actions through graphical user interfaces, raising safety concerns in real-world deployment. Existing safety benchmarks require substantial manual effort to expand, often overlook unsafe responses, such as unsafe intentions or failed unsafe attempts, that do not lead to harmful final outcomes, and provide limited evidence for attributing unsafe responses to specific risk factors, which hinders extending test cases to new sce- narios. To address these limitations, we introduce AlphaCUA, an automated red- teaming framework for scalable and interpretable safety evaluation of CUAs, built on three collaborating agents. A generation agent, guided by reusable skills that encode quality standards and safety objectives, constructs executable safety tasks. An analysis agent examines every execution trajectory regardless of its final out- come, attributing observed unsafe responses to candidate risk factors for success- ful attacks and inferring candidate factors from the agent’s observed actions and expressed intentions for failed ones. A validation agent tests every candidate fac- tor, attributed or inferred, by constructing task variants that differ solely in that fac- tor and checking whether the target CUA’s response changes accordingly; inferred factors showing partial progress are refined and retested until they are validated or a refinement limit is reached. Validated factors are fed back to the generation agent to construct tasks in new scenarios, closing the loop between risk discovery and benchmark expansion. Applying AlphaCUA, we construct ALPHABENCH-SEED from initially generated tasks and ALPHABENCH-TRANSFERRED from validated risk factors, each with 100 tasks spanning multiple applications and threat types. Of these 200 tasks, 96.5% pass human quality review without revision and the rest need only minor corrections, so benchmark expansion requires minimal human effort. Evaluating four CUAs on ALPHABENCH-SEED, we find that trajectory- based attack success rate (ASR) exceeds final-state ASR by 8.75 percentage points on average, exposing unsafe responses that outcome-based evaluation misses. ALPHABENCH-TRANSFERRED, which instantiates validated risk factors in new scenarios, raises the mean trajectory-based ASR from 44.75% to 60.50% with consistent gains across all four models, indicating that the discovered factors are both effective and transferable across scenarios. Our code is available at https://anonymous.4open.science/r/AlphaCUA_2027-7CBD/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.