acceptodds
Under review as a conference paper at ICLR 2027

The Entropy Bias of Hybrid Attention: Training Models to Be Sparse-Ready

Abstract

Hybrid language models improve long-context efficiency by concentrating global information exchange into a small number of full-attention layers. We find that these layers exhibit a systematic entropy bias: compared with full-attention Transformers, their attention is more diffuse, leaving less mass recoverable under a fixed sparse budget and making them less robust to sparsification. Importantly, this behavior is not an inherent limitation of hybrid architectures, but an optimization outcome that can be shaped during training. Through controlled pretraining experiments, we show that an architecture-aware concentration prior reduces entropy and substantially improves sparse long-context performance. The gains extend across hybrid architectures, model sizes, and longer contexts, and remain evident with learned QK gains and trained sparse indexers. Entropy saturation further offers a practical guide to prior strength. Together, these findings establish architecture-aware attention concentration as a key training principle for sparse-ready hybrid models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.