acceptodds
Under review as a conference paper at ICLR 2027

VirtualFelonyBench: Do LLM Agents Hack to Complete Impossible Tasks?

Abstract

Agent safety research has mostly studied harm that starts from a malicious input. Incidents in 2026 show that deployed agents with benign instructions acted beyond their permissions without being asked to and rarely reported doing so. We call such behavior an agent-initiated authorization violation. We present VirtualFelonyBench, a benchmark of 180 base tasks in self-hosted deployments of real applications. Each base task gives an agent an ordinary goal that its account is not permitted to reach. Every base task has four matched conditions. Three of them contain an injected vulnerability that makes the goal reachable and differ in what the workspace says about it: nothing, a formal maintenance record, or a note from a peer agent. We evaluate three deployed agent systems in 2,160 trials. When nothing describes the vulnerability, the controller verifies no exploitation, but the systems differ sharply in how often they stop and report that the task cannot be completed. Once the vulnerability is described, at least one system exploits it on 160 of the 180 base tasks. A formal record also leads to exploitation, and neither form of description is consistently stronger. In 281 of all 451 exploitations (62%), the agent reports an ordinary success or gives no reply. Safety evaluations must therefore measure separately whether an agent exploits a vulnerability, stops honestly, and discloses what it did.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.