LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization
Abstract
Direct Preference Optimization (DPO) has become a standard reward-model-free approach for aligning language models with preference data. However, as the scaled preference margin grows during training, the logistic DPO loss becomes progressively less sensitive to further changes. We study DPO from a loss-level geometric perspective and identify the sigmoid factor as a learning signal that characterizes the local sensitivity of the objective. Based on this view, we propose Learning-Signal-Controlled Direct Preference Optimization (LSC-DPO), which dynamically regulates the learning signal near a target regime. A log-space analysis establishes conditions for stable tracking of the target learning-signal regime. Experiments on AlpacaEval 2, MT-Bench, and Anthropic-HH show that LSC-DPO consistently improves over DPO and strong preference-optimization baselines. We further observe that learning-signal dynamics are systematically related to the initial coefficient. We then investigate this relationship and derive a signal-budget compensation rule that substantially reduces performance variation across coefficient initializations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.