acceptodds
Under review as a conference paper at ICLR 2027

HackFree-GRPO: Detecting and Mitigating Reward Hacking in RLVR- Trained Code Agents

Abstract

Reinforcement learning with verifiable rewards can reward code agents for passing tests without resolving the underlying issue. We study this failure mode in repository-level, multi-turn training and introduce HackFree-GRPO. Its frozen detector scores actions, patch diffs, and execution logs. Reward shaping discounts risky successes, and trajectory filtering removes high-risk samples before group-relative advantages are computed. Shaping alone can leave an exploit positively advantaged in a low-success group; after filtering, score differences can still distinguish risky survivors. On SWE-bench Verified, FULL reduces protocol-defined hacking from 41.7% to 5.8% and raises strict-audit pass@1 from 30.5% to 49.0%, while verifier acceptance is 52.0% versus 52.3%. Adversarial-test-only pass@1 increases from 31.2% to 49.9%, exceeding SHAPE by 2.5 percentage points. We derive the exact change in retained-group advantages and identify the conditions under which filtering removes residual positive reinforcement. The analysis specifies a group-level mechanism that can be tested alongside the observed task-level gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.