acceptodds
Under review as a conference paper at ICLR 2027

Auditing Reward Hacking for Agent Environments via World Modeling

Abstract

Reward hacking allows language-model agents to earn high benchmark scores without completing the intended task. Existing approaches often detect hacking only after an agent has finished, or search for particular attack patterns whose coverage of other exploits remains uncertain. We introduce RewardProof, an agent-driven world-modeling framework for proactive auditing of agent environments. Its auditing agent explores and simulates potential hacking worlds, such as a state containing a forged reward file, and validates their effects by running submissions with and without the candidate attack before deployment. We introduce a fine-grained auditing protocol that distinguishes validated reward effects, unsuccessful attempts targeting the verifier, other unvalidated proposals, and cases with insufficient evidence. Across 567 environments from 18 benchmark families, we validate 1,116 attacks in 79.9% of environments and identify 44 environments whose scorers do not distinguish the tested submissions. Seventeen attacks also reproduce in the upstream SWE-bench Verified harness. On 72 shared tasks, our proposal method finds more validated attacks at higher precision than the evaluated proposal baselines. These results show how proactive auditing can give benchmark developers concrete evidence of scoring shortcuts before deploying an evaluation environment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.