Less Is More in DPO: Learning from the Lowest-Entropy 20% of Tokens
Abstract
Direct Preference Optimization (DPO) is widely used to align large language models with human preferences. However, its sequence-level objective does not explicitly distinguish the contributions of individual tokens. We find that optimizing only a small subset of low-entropy tokens can outperform using all response tokens. Based on this observation, we propose **Low-Entropy Support DPO (LES-DPO)**. LES-DPO ranks tokens by the policy model's predictive entropy. It selects a low-entropy subset separately from each chosen and rejected response. The preference loss uses only these selected positions. LES-DPO requires no auxiliary models, additional policy forward passes, or extra trainable parameters. Experiments on six Qwen2.5, Llama-3.1, and Mistral models show consistent improvements on AlpacaEval 2, Arena-Hard, and MT-Bench. Gradient analysis shows that high-entropy tokens tend to produce larger gradients. Their directions are less consistent and conflict more often. These findings suggest that focusing on low-entropy tokens can provide more stable signals for preference learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.