How Far Can Sliding Window Attention See?
Abstract
Sliding window attention (SWA) reduces the inference cost of Transformers by restricting each token to attend only to a fixed-size local window. Although the resulting efficiency gains are well understood, it remains unclear how information propagates beyond this window and, consequently, the maximal sentence length a SWA model can effectively manipulate. We provide an intuitive account of information propagation through successive SWA layers and use it to derive a theoretical prediction for the maximal effective length. For a model with window size and layers, our analysis predicts a context length of , where is the -th harmonic number. Experiments on the cramming task across several model families, scales, depths, and window sizes confirm this prediction, with predicted lengths within of the measured limits in all tested configurations. Finally, we give an explicit construction that approximates full attention with SWA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.