acceptodds
Under review as a conference paper at ICLR 2027

Likelihood Ratio Aware Data Selection for Direct Preference Learning

Abstract

Direct Preference Optimization (DPO) is a widely used offline preference alignment method for large language models (LLMs), and its performance is highly dependent on the quality of the preference data. Conventional selection methods prioritize the separability between paired responses, but we argue that this overlooks the inherent value of individual samples related to the true data distribution. To address this, we propose a likelihood ratio aware data selection framework that identifies high value samples by integrating distributional reliability with data separability. Our method enhances training stability by filtering out low likelihood ratio regions whose gradients are large and directionally incoherent, while achieving superior alignment performance even with a small subset of data that surpasses full dataset training. Extensive experiments across multiple large language models demonstrate that our approach consistently outperforms standard DPO, its variants, and existing selection baselines, and generalizes well to other preference learning losses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.