The Entropy Bias of Hybrid Attention: Training Models to Be Sparse-Ready
Abstract
Hybrid language models improve long-context efficiency by concentrating global information exchange into a small number of full-attention layers. We find that these layers exhibit a systematic entropy bias: compared with full-attention Transformers, their attention is more diffuse, leaving less mass recoverable under a fixed sparse budget and making them less robust to sparsification. Importantly, this behavior is not an inherent limitation of hybrid architectures, but an optimization outcome that can be shaped during training. Through controlled pretraining experiments, we show that an architecture-aware concentration prior reduces entropy and substantially improves sparse long-context performance. The gains extend across hybrid architectures, model sizes, and longer contexts, and remain evident with learned QK gains and trained sparse indexers. Entropy saturation further offers a practical guide to prior strength. Together, these findings establish architecture-aware attention concentration as a key training principle for sparse-ready hybrid models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.