acceptodds
Under review as a conference paper at ICLR 2027

Less Drift, Better Alignment: Selecting the Fittest Preference Data for Offline DPO

Abstract

Offline Direct Preference Optimization (DPO) relies on pre-collected preference data, but mismatches between such data and the current policy can cause large local perturbations, ultimately leading to excessive policy drift. To address this issue, we seek to identify preference pairs that induce smaller local perturbations to the current policy when trained on. We study computable proxies for this local update cost and identify a generation-level asymmetry between chosen and rejected responses. Controlled generation and gradient analyses show that chosen-side updates affect decoding more directly and reliably, whereas rejected-side effects depend on their overlap with the current policy's high-probability generation region. The chosen-gradient norm therefore provides a better proxy for local update cost. We further show that this cost is only weakly coupled with preference quality and semantic coverage, so low-cost subsets need not sacrifice either. Motivated by these results, we propose Low-, which ranks complete preference pairs by a length-normalized chosen-gradient score and selects the lowest-scoring subset under a fixed data budget before standard DPO training. Under local smoothness, lower cost tightens the upper bound on log-probability perturbation and raises the lower bound on retained probability mass for high-quality responses in the generation region. Across four models, three datasets, and five metrics, Low- ranks in the top two in 55 of 60 settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.