acceptodds
Under review as a conference paper at ICLR 2027

Mixture of Convolutions: Token-Adaptive Short Convolution for Sequence Modeling

Abstract

Short convolution has become a common local mixing component in modern sequence models, yet existing variants are typically static: a single learned kernel defines the same local aggregation rule at every position. We introduce Mixture of Convolutions (MoC), a simple adaptive generalization of short convolution that replaces this single kernel with a mixture over learned convolutional bases. Evaluated in both Transformer and Gated DeltaNet architectures, MoC shows broad improvements in language-modeling perplexity over static short convolution, at modest additional cost. We additionally study how short convolutions are optimized, and find that their performance is sensitive to the interaction between initialization scale and learning rate, in a way that affects static short convolution as well as MoC. These results suggest that token-adaptive local mixing is a simple and effective extension of a widely used sequence-modeling primitive.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.