Less Drift, Better Alignment: Selecting the Fittest Preference Data for Offline DPO
Abstract
Offline Direct Preference Optimization (DPO) relies on pre-collected preference data, but mismatches between such data and the current policy can cause large local perturbations, ultimately leading to excessive policy drift. To address this issue, we seek to identify preference pairs that induce smaller local perturbations to the current policy when trained on. We study computable proxies for this local update cost and identify a generation-level asymmetry between chosen and rejected responses. Controlled generation and gradient analyses show that chosen-side updates affect decoding more directly and reliably, whereas rejected-side effects depend on their overlap with the current policy's high-probability generation region. The chosen-gradient norm therefore provides a better proxy for local update cost. We further show that this cost is only weakly coupled with preference quality and semantic coverage, so low-cost subsets need not sacrifice either. Motivated by these results, we propose Low-, which ranks complete preference pairs by a length-normalized chosen-gradient score and selects the lowest-scoring subset under a fixed data budget before standard DPO training. Under local smoothness, lower cost tightens the upper bound on log-probability perturbation and raises the lower bound on retained probability mass for high-quality responses in the generation region. Across four models, three datasets, and five metrics, Low- ranks in the top two in 55 of 60 settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.