acceptodds
Under review as a conference paper at ICLR 2027

Oberon: Does Detecting Reward Hacking Inform Its Removal?

Abstract

Reward models (RMs) are central to language model alignment, and interventions on their internal representations can mitigate reward biases. However, a direction that reliably detects a response property does not necessarily reduce its associated reward gap when erased. To address this challenge, we introduce Oberon, which jointly learns a single direction for detection and reward-gap attenuation in a frozen RM. We first formalize directional readout and representation erasure within an observation–intervention framework and show that detector scores alone do not generally identify signed erasure effects over a family of affine downstream maps. We then derive a pairwise upper bound on a population risk combining both objectives and use its empirical form to couple a paired logistic loss with a penalty on reward gaps remaining after erasure of the same direction. We evaluate five RMs on fixed response pairs involving formatting, length, sycophancy, and conflicts between quality and style. In each condition, Oberon achieves the highest average proportion of pairs with both correct paired detection and reduced absolute reward gaps among the compared methods, with an average drop in preference accuracy below one percentage point. On formatting and length, the joint success rates averaged across the five RMs reach 90.5% and 70.3%, compared with 83.5% for score penalty and 64.0% for difference of means, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.