HackBench: A Benchmarking Suite for Evaluating Naturally-Emerging Reward Hacking in LLMs
Abstract
Reward hacking–when an LLM agent satisfies a proxy reward signal without achieving the actual underlying objective–is a central integrity risk for deploying agents in autonomous settings. However, despite being a widely recognized concern, existing LLM reward hacking benchmarks focus predominantly on coding tasks. To bridge this gap, we present HackBenchCode: https://anonymous.4open.science/r/HACKBENCH-RELEASE-1194, a novel benchmarking suite that aims to manifest reward hacking behavior through naturalistic and stylized game-based scenarios. Specifically, HackBench integrates three existing benchmarks of agentic workflows (DBBench, TAU-Bench, and WorkBench) with two game-based benchmarks (maze and grid), and transforms them into reward hacking benchmarks by introducing exploits that provide “illegal” shortcuts to solving tasks. Moreover, our use of game-based scenarios further enables controlled difficulty across three LLM interaction levels: one-shot, interactive (with tool access), and fully autonomous shell-based deployment. We illustrate our benchmark across several frontier LLM agents and show that reward hacking is strongly condition- and domain-dependent, and rises sharply under pressure and competition, especially when tasks are otherwise unsolvable. Across over 5,000 runs under a shared reward-hack taxonomy, we report results on two metrics: a detected hack rate and a deliberate hack rate (for which we confirm intent based on reasoning traces). Because every task has a ground-truth evaluation criterion, detection is deterministic and judge-free, based on file-tampering checks and exploit-tool calls. Finally, we surface an important hacking-considered-only signal where agents articulate cheat plans without acting on them—a behavior invisible to our main metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.