acceptodds
Under review as a conference paper at ICLR 2027

How Far Can Sliding Window Attention See?

Abstract

Sliding window attention (SWA) reduces the inference cost of Transformers by restricting each token to attend only to a fixed-size local window. Although the resulting efficiency gains are well understood, it remains unclear how information propagates beyond this window and, consequently, the maximal sentence length a SWA model can effectively manipulate. We provide an intuitive account of information propagation through successive SWA layers and use it to derive a theoretical prediction for the maximal effective length. For a model with window size and layers, our analysis predicts a context length of , where is the -th harmonic number. Experiments on the cramming task across several model families, scales, depths, and window sizes confirm this prediction, with predicted lengths within of the measured limits in all tested configurations. Finally, we give an explicit construction that approximates full attention with SWA.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.