Likelihood Ratio Aware Data Selection for Direct Preference Learning
Abstract
Direct Preference Optimization (DPO) is a widely used offline preference alignment method for large language models (LLMs), and its performance is highly dependent on the quality of the preference data. Conventional selection methods prioritize the separability between paired responses, but we argue that this overlooks the inherent value of individual samples related to the true data distribution. To address this, we propose a likelihood ratio aware data selection framework that identifies high value samples by integrating distributional reliability with data separability. Our method enhances training stability by filtering out low likelihood ratio regions whose gradients are large and directionally incoherent, while achieving superior alignment performance even with a small subset of data that surpasses full dataset training. Extensive experiments across multiple large language models demonstrate that our approach consistently outperforms standard DPO, its variants, and existing selection baselines, and generalizes well to other preference learning losses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.