Auditing Reward Hacking for Agent Environments via World Modeling
Abstract
Reward hacking allows language-model agents to earn high benchmark scores without completing the intended task. Existing approaches often detect hacking only after an agent has finished, or search for particular attack patterns whose coverage of other exploits remains uncertain. We introduce RewardProof, an agent-driven world-modeling framework for proactive auditing of agent environments. Its auditing agent explores and simulates potential hacking worlds, such as a state containing a forged reward file, and validates their effects by running submissions with and without the candidate attack before deployment. We introduce a fine-grained auditing protocol that distinguishes validated reward effects, unsuccessful attempts targeting the verifier, other unvalidated proposals, and cases with insufficient evidence. Across 567 environments from 18 benchmark families, we validate 1,116 attacks in 79.9% of environments and identify 44 environments whose scorers do not distinguish the tested submissions. Seventeen attacks also reproduce in the upstream SWE-bench Verified harness. On 72 shared tasks, our proposal method finds more validated attacks at higher precision than the evaluated proposal baselines. These results show how proactive auditing can give benchmark developers concrete evidence of scoring shortcuts before deploying an evaluation environment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.