acceptodds
Under review as a conference paper at ICLR 2027

Information-Aligned Preference Optimization for Direct Preference Learning

Abstract

Direct Preference Optimization (DPO) has become a standard approach for aligning large language models with human feedback. Recent work has addressed its sensitivity to response length, yet equalizing token counts alone leaves open which tokens should contribute to preference learning. In this work, we study this question through an information-theoretic lens and introduce Information-Aligned Preference Optimization (IAPO). IAPO uses model entropy as a practical ranking signal, masking low-entropy positions in the longer response to match the retained token count of the shorter response. Our analysis shows that true conditional entropy upper-bounds token-level preference information and that top-K selection minimizes a discarded-entropy surrogate under a fixed token budget. Experiments across multiple model families and scales show that IAPO improves over DPO and achieves strong average performance against recent variants. Matched-budget ablations further demonstrate the benefit of entropy-guided selection beyond length equalization, supporting selective token aggregation as an effective strategy for preference optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.