acceptodds
Under review as a conference paper at ICLR 2027

Not All Tokens Are Equal: Entropy-Guided Credit Assignment for Direct Preference Optimization

Abstract

Direct Preference Optimization (DPO) aligns language models with a single sequence-level preference signal, leaving open how that signal should be attributed to individual tokens. We show empirically that this attribution is highly uneven: under the current policy, a small minority of high-entropy positions accounts for most of the implicit reward margin, whereas the low-entropy majority contributes almost nothing. Motivated by this observation, we propose Entropy-Guided DPO (EG-DPO), which reweights token-level policy–reference log-ratios by a detached, monotone function of the policy's own predictive entropy. The weights are normalized within each response, so the objective coincides with DPO when they are uniform, and they are recomputed at every step, so they track the evolving policy without adding a gradient path, an auxiliary model, or a forward pass. We justify the design with two complementary analyses: a signal-to-noise argument showing that uniform aggregation dilutes preference-relevant gradients, and a bound showing that the attainable pairwise preference-divergence at a position grows with the policy's entropy there. Across Llama-3-8B, Mistral-7B, and four Qwen2.5-Instruct scales, EG-DPO improves over DPO and recent token-level baselines on AlpacaEval 2, Arena-Hard v0.1, and MT-Bench, and yields the best average preference-classification accuracy on HH-RLHF, with its largest gain at the 14B scale.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.