Act on Your Behalf, Beyond Your Authorization: Benchmarking LLM Agent Authorization Alignment
Abstract
When an LLM agent acts on behalf of a user, it is expected to remain within the authorization scope the user grants. Existing agent evaluations largely score task completion or harmful consequences and can miss actions that violate this delegation without producing conventional safety or privacy harms. We present AuthorizationBench, a benchmark of authorization alignment grounded in 44 real-world out-of-authorization cases. A scalable construction pipeline produces 286 executable tasks spanning situational triggers, direct harmful consequences, and ten domains. We evaluate four models at two reasoning-effort settings with three harnesses, producing 6,864 trajectories in an authorization-aware sandbox. The out-of-authorization rate (OAR) ranges from 10.5% to 26.6% across the 24 configurations and is 19.2% overall. Task completion does not imply authorization, and higher reasoning effort does not reliably reduce OAR. Agents act out of authorization at similar rates with and without direct harmful consequences, while benign pressures elicit more OA behavior than malicious injections (23.5% vs 17.6%, ). Agents that ask the principal before completing a task have a lower OAR, although the relationship is correlational. Our results show that task capability, additional reasoning, and consequence-oriented safety evaluation do not by themselves ensure authorization alignment, motivating it as a distinct target for agent evaluation and development.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.