Structure Before Attention: A Zero-Initialized Convolution Recovers Windowed Language-Model Quality
Abstract
Sliding-window attention reduces long-context language-modeling cost but can worsen perplexity. We study whether a small convolution added to the residual stream before attention can offset this increase, and what contributes to its benefit. Under paired-seed attribution, paired runs share initialization, data order, and evaluation batches, differing only by the initially inactive branch. Under dense attention the branch gives modest gains at 81M and 253M parameters. In 81M models on PG-19 book text with an 8192-token context and a 256-token window in eight of ten layers, it lowers perplexity by 2.1%, against 0.6% under dense attention, and the windowed model with the branch finishes within 0.16% of the dense model with it in each of three seeds. Across contexts and window sizes, the extra gain descriptively tracks the window penalty (the baseline cost of using sliding-window attention). A pointwise (width-one) branch, which mixes no neighboring tokens, recovers most of the gain, and in a single-seed test, removing the branch after 60% of training and training on retains most of its advantage. The intuition that dense attention learns the branch's function while windowed attention cannot motivated an amplitude-sensitivity test whose registered predictions fail; the mechanism behind the amplified benefit remains open.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.