acceptodds
Under review as a conference paper at ICLR 2027

Disentangling Optimization Scale from Preference Scale in DPO

Abstract

Direct Preference Optimization (DPO) uses β both to control the sharpness of the preference likelihood and, through an explicit gradient prefactor, to rescale optimization. Under fixed-learning-rate, finite-budget training, this coupling can make policy movement non-monotone in β: our sweeps exhibit near-zero KL movement at very small β (validation KL below 0.05 nats per token), a peak at an intermediate value, and reduced movement at larger values. A scalar stopping-tolerance model explains this qualitative pattern. We also show that standard DPO losses are difficult to compare across β, because nearly identical loss curves can correspond to substantially different margins and KL divergences. We propose a normalized centered-softplus objective that is a positive affine transformation of DPO for every fixed β > 0, preserving its ordering and any attained minimizers while removing the explicit β gradient prefactor. This parameterization separates likelihood sharpness from first-order update scale and yields more informative loss trajectories in our experiments. It also has a continuous β → 0 limit corresponding to a linear preference-margin objective.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.