acceptodds
Under review as a conference paper at ICLR 2027

Raising the Floor, Lowering the Ceiling: Adaptive Reward Bounds for Direct Preference Optimization

Abstract

An increase in the preference margin optimized by Direct Preference Optimization (DPO) does not imply an increase in the probability assigned to preferred behavior. The margin can also increase if the chosen and rejected rewards both decrease, provided that the rejected decreases more rapidly. This behavior reflects an under-specification in the DPO objective. It constrains only the difference between the policy's and reference model's response scores, rather than their individual evolution. Moreover, the shared prompt, token representations, and parameters induce coupled gradients for the two responses, allowing a simultaneous decrease to constitute a favorable descent direction for the pairwise objective. We propose Adaptive Reward Bounds (ARB), a trajectory-aware regularization framework that augments DPO with explicit directional control. At each step, ARB computes length-normalized changes in the reward, aggregates chosen and rejected values over the batch, and tracks them with sliding-windows. It imposes a lower bound on the chosen statistic and an upper bound on the rejected statistic using detached soft penalties only upon violations. Different from existing methods, ARB adapts its constraints to the training trajectory, tolerates local reversals, and preserves directional pressure through optional biases. Experiments confirm that ARB effectively mitigates the undesired joint decline of chosen and rejected rewards while preserving the preference learning objective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.