Information-Aligned Preference Optimization for Direct Preference Learning
Abstract
Direct Preference Optimization (DPO) has become a standard approach for aligning large language models with human feedback. Recent work has addressed its sensitivity to response length, yet equalizing token counts alone leaves open which tokens should contribute to preference learning. In this work, we study this question through an information-theoretic lens and introduce Information-Aligned Preference Optimization (IAPO). IAPO uses model entropy as a practical ranking signal, masking low-entropy positions in the longer response to match the retained token count of the shorter response. Our analysis shows that true conditional entropy upper-bounds token-level preference information and that top-K selection minimizes a discarded-entropy surrogate under a fixed token budget. Experiments across multiple model families and scales show that IAPO improves over DPO and achieves strong average performance against recent variants. Matched-budget ablations further demonstrate the benefit of entropy-guided selection beyond length equalization, supporting selective token aggregation as an effective strategy for preference optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.