acceptodds
Under review as a conference paper at ICLR 2027

DimPO: Dimensionality Reduction for Attention using Preference Optimization

Abstract

A linear projection might reduce the dimension of the query and key vectors in attention without updating the pretrained model, but it remains unclear which training objective best preserves the behavior of that model. The straightforward objective is to match the full attention distribution, for example by minimizing the KL divergence. We instead ask whether a preference over keys relative to a query, and preserving the attention mass of only the highest-weighted keys, provides a better training signal, especially in long-context settings. Existing reference-free preference objectives are mostly pairwise and underuse the full key set that listwise and KL supervision already use. We therefore introduce DimPO, a loss specific to this projection, which combines listwise preference optimization with no reference model and a lightweight top-k cross-entropy term for head-fidelity supervision. To verify the direct effect of our loss on how faithfully attention is preserved, DimPO is trained offline from the attention patterns of a frozen language model, with the query and key projections optimized independently per layer. Across LLaMA3.2-3B, LLaMA3.1-8B, Qwen2.5-7B, and Qwen3-4B Instruct models, pairwise preference objectives outperform the triplet objective baseline and, on short-context tasks, retain 98% of the original score when a projection to half the dimension is applied to at most the last 40% of the attention layers. When applied to more layers or evaluated on long-context RULER, however, they no longer maintain a reasonable performance. In contrast, KL and DimPO, which use every key during training, still retain about 95% of the original RULER 4k score on the 8B model when applied to up to 50% of the layers. KL-based projections nevertheless remain closer to the original attention distribution in terms of KL divergence and closer in the MSE of the attention output, yet DimPO achieves better downstream performance. The difference becomes increasingly pronounced once the projection covers more than 50% of the attention layers, where preserving preferences over keys relative to the query, together with head-fidelity on the highest-weighted keys, improves performance on more complex tasks, including SQuAD, common-word extraction, frequent-word extraction, and variable tracking. These results suggest that, under a reduced query and key dimension, preserving the ordering and concentration of task-relevant attention can matter more than minimizing the discrepancy of the full distribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.