acceptodds
Under review as a conference paper at ICLR 2027

LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization

Abstract

Direct Preference Optimization (DPO) has become a standard reward-model-free approach for aligning language models with preference data. However, as the scaled preference margin grows during training, the logistic DPO loss becomes progressively less sensitive to further changes. We study DPO from a loss-level geometric perspective and identify the sigmoid factor as a learning signal that characterizes the local sensitivity of the objective. Based on this view, we propose Learning-Signal-Controlled Direct Preference Optimization (LSC-DPO), which dynamically regulates the learning signal near a target regime. A log-space analysis establishes conditions for stable tracking of the target learning-signal regime. Experiments on AlpacaEval 2, MT-Bench, and Anthropic-HH show that LSC-DPO consistently improves over DPO and strong preference-optimization baselines. We further observe that learning-signal dynamics are systematically related to the initial coefficient. We then investigate this relationship and derive a signal-budget compensation rule that substantially reduces performance variation across coefficient initializations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.