When Shorter Is Better: Very Short Sliding Windows in Gated Hybrid Attention
Abstract
As modern applications demand larger models and longer contexts, global attention has become a dominant cost in both training and inference. Hybrid attention that combine sliding-window attention with global attention reduce this cost, but are typically viewed as trading quality for efficiency, with larger windows assumed to be better. In this paper, we revisit this view through the lens of recent findings on pretrained LLMs, in particular the literature on attention sinks. We find that hybrid architectures with very short sliding windows (16 or 32 tokens) outperform both the full global attention baseline and hybrids with larger windows, even though the latter are strictly more expressive. Adding conditional gating that dynamically balances head contributions further amplifies these gains. We validate both findings with ablations over head configurations, architectural variants, sequence lengths, and training budgets. Our results challenge the conventional trade-off between window size and quality, suggesting that short windows act as a beneficial inductive bias rather than a cheap approximation to global attention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.