acceptodds
Under review as a conference paper at ICLR 2027

LCPO: Towards Addressing Reward Degeneration in Direct Preference Optimization

Abstract

Direct Preference Optimization (DPO) and its variants are widely used for preference alignment in large language models. However, prior work has shown that DPO often suffers from reward degeneration during training, where the implicit rewards of both chosen and rejected responses drift into extremely negative regions. In this paper, we analyze this phenomenon through the lens of reward margin and reward level. We show that DPO enlarges the reward margin but provides no lower-bound control along the reward-level direction. Consequently, once optimization enters a low-level region, vanilla DPO has no intrinsic mechanism to restore the reward level. We further show that pairwise preference data do not identify the absolute reward value, so correcting reward degeneration necessarily requires introducing additional priors. This motivates us to formulate degeneration correction as a prior design problem. Under this view, we propose three design principles for such priors: level-active, target-preserving, and calibration-free. Guided by these principles, we introduce Level-Constrained Preference Optimization (LCPO), a simple and general solution that imposes a relational constraint between chosen and rejected rewards to counteract reward-level collapse. Empirical results on representative benchmarks show that LCPO effectively mitigates reward degeneration and consistently outperforms existing baselines. Code is provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.