acceptodds
Under review as a conference paper at ICLR 2027

Reward Hacking Reveals Itself in Judge Confidence

Abstract

Reinforcement learning with verifiable rewards has driven major advances in reasoning. Yet reliable rule-based verification remains challenging for many tasks, motivating the use of LLM judges as scalable proxy reward providers. However, optimizing against an LLM judge can induce reward hacking, where policies exploit its evaluation errors to increase proxy reward while task performance stagnates or deteriorates. Existing approaches to monitoring reward hacking often require external evaluation, additional inference, or access to model internals. In contrast, we uncover a monitoring signal already present in the LLM judge's outputs: verdict confidence systematically declines as reward hacking develops. Building on this finding, we propose **balanced verdict margin (BVM)**, a lightweight and task-agnostic metric that tracks reward hacking using the judge's verdict-token log-probabilities. We further develop a **margin-guided reward intervention (MRI)** that mitigates reward hacking without requiring gold or reference feedback during training. Extensive experiments show that BVM consistently tracks reward hacking across tasks, while MRI substantially improves task performance over unmodified LLM-judge training, even matching the rule-based control in mathematical reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.