LCPO: Towards Addressing Reward Degeneration in Direct Preference Optimization
Abstract
Direct Preference Optimization (DPO) and its variants are widely used for preference alignment in large language models. However, prior work has shown that DPO often suffers from reward degeneration during training, where the implicit rewards of both chosen and rejected responses drift into extremely negative regions. In this paper, we analyze this phenomenon through the lens of reward margin and reward level. We show that DPO enlarges the reward margin but provides no lower-bound control along the reward-level direction. Consequently, once optimization enters a low-level region, vanilla DPO has no intrinsic mechanism to restore the reward level. We further show that pairwise preference data do not identify the absolute reward value, so correcting reward degeneration necessarily requires introducing additional priors. This motivates us to formulate degeneration correction as a prior design problem. Under this view, we propose three design principles for such priors: level-active, target-preserving, and calibration-free. Guided by these principles, we introduce Level-Constrained Preference Optimization (LCPO), a simple and general solution that imposes a relational constraint between chosen and rejected rewards to counteract reward-level collapse. Empirical results on representative benchmarks show that LCPO effectively mitigates reward degeneration and consistently outperforms existing baselines. Code is provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.