Efficient Reward Hacking Detection using Answer Likelihood under Truncated Reasoning
Abstract
Reinforcement learning (RL) for language models is vulnerable to reward hacking, in which the model obtains high rewards by exploiting flaws in the data or reward signal rather than correctly solving the intended task. Chain-of-Thought (CoT) monitoring can flag such hacking that the model reveals in its reasoning, but misses implicit reward hacking, where the CoT appears benign. TRACE (Wang et al., 2026) addresses this problem by checking whether answers regenerated from truncated CoTs already earn high reward, but requires generation and reward evaluation at every cutoff. We introduce (n-K ikelihood valuation), which instead scores the likelihood of the model's own final answer under CoT truncation. At each cutoff, MILE teacher-forces this answer to the model, normalizes each token's log-probability against the corresponding next-token distribution, and averages over the lowest-scoring % of tokens. Integrating these scores over the retained CoT fraction yields the detection score, requiring no additional decoding or reward evaluation once the full response has been initially generated. We evaluate MILE on math and code tasks under three controlled loopholes: answer hints in the prompt, a misspecified reward rule, and memorization of training problems. Across models and loophole types, MILE achieves mean F1 and AUROC comparable to TRACE while scoring - faster, and consistently outperforms CoT monitoring even with a larger model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.