acceptodds
Under review as a conference paper at ICLR 2027

The Many Faces of Short Convolutions in Transformers

Abstract

Short causal convolutions improve the performance of sequence models, a benefit often attributed to the formation of single-layer induction heads. By disentangling the scaling and shifting operations within a convolution, we investigate their role from a **stability** perspective, showing that their benefits extend beyond expressivity. Starting from a simple channel-wise gain, we discover that convolutions mitigate rank-collapse by functioning as learnable **residual scaling**. However, this is not the only benefit for trainability. By extending the filter size, we show that convolutions alter the **positional information** learned through causal masks (NoPE), improving representational diversity. This enables training an otherwise unstable post-norm Transformers, while broadening the range of viable learning rates for pre-norm Transformers. Together, these findings identify channel-wise scaling and local temporal mixing as complementary contributors to the effectiveness of short convolutions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.