acceptodds
Under review as a conference paper at ICLR 2027

AUTHBENCH: Benchmarking Interactive User Authorization in Computer-Use Agents

Abstract

Computer-use agents act on their users' behalf, and delegation is legitimate only if every consequential action stays within what the user permitted. However, existing benchmarks ask only whether the agent reached the requested state, and those that score constraints fix them in advance, so nothing said during a run can grant or withdraw a permission. We argue that authorization is instead an in-run grant: a permission the user issues while the agent works, bound to the values they were shown and void once those values change. Grants are private, perishable, and invisible in the final state. We realize this in AUTHBENCH, an interactive environment over complete web applications, where a simulated user holds authorization privately, state can change underneath an approved action, and an append-only trace records every request and attempted action. On it we build a benchmark of authorization-sensitive tasks with executable oracles and evaluate three deployed computer-use systems. Our results show that capability and authorization come apart: the system that completes the most tasks stays within its user's grants least often, none is both successful and authorized in even half its runs, and not one confirmation was bound to the action it authorized. AUTHBENCH thus makes authorization measurable as a property distinct from capability, and we release the environment, the tasks, and every recorded run.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.