Stop Training on Pairs You've Already Learned: Policy-Aware Pair Selection for Preference Optimization
Abstract
Reinforcement learning from human feedback (RLHF) is the standard approach to aligning language models with human preferences, and in practice it is most often performed offline with direct methods such as DPO that learn from preference pairs alone. What such methods learn depends heavily on which pairs they see. Following the Delta Learning Hypothesis, which pairs responses from a large and a small model from the same model class, a strong recent rule (MaxMin) selects the best and worst responses from a candidate pool for each prompt. We show that DPO with MaxMin data selection saturates the training signal quickly. In particular, the policy quickly separates the best from the worst responses, its implicit reward margins grow, and the objective's sigmoid flattens, resulting in a small gradient norm and weight updates. We propose ActiveDelta, an adaptive selection rule that follows the policy being trained: for each prompt, it picks the largest-quality-gap pair according to an external oracle feedback which the current policy has not yet confidently separated, keeping gaps large while avoiding saturation. ActiveDelta consistently outperforms tuned MaxMin selection on downstream benchmarks and reward-model scores on a held-out set of prompts. Crucially, ActiveDelta keeps improving where MaxMin plateaus; on the benchmark suite it matches MaxMin with one-sixth of the prompts. A practical variant that ranks candidates with a learned reward model instead of an oracle and queries the reward oracle only for the selected pair retains these gains, which makes ActiveDelta applicable when scoring every response is too costly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.