acceptodds
Under review as a conference paper at ICLR 2027

Spectral Polynomial Linear Attention (SPLA): polynomial singular-value maps for gated linear models

Abstract

Linear models have gained widespread attention as a viable alternative to reduce the cost of full-attention Transformers. However, due to their fixed-size state space, the spectral behavior of the state space directly determines their performance. Recent work has proposed that their dominant components cause the effective rank of the state to collapse, making it difficult to maintain long-range memory capacity. We draw on the ideas of the Muon optimizer and focus on recent work such as MesaNet, finding that it essentially regulates the key-value matrix of all past states. We argue that a broader singular value spectrum of the keys and values may provide the model with more choices. Based on this, we propose Spectral Polynomial Linear Attention (SPLA). The architecture we propose provides an alternative way of using the state of linear models through high-order polynomial approximation function transformations. We further conduct experiments on different tasks, and the results show that our method effectively preserves the long-term recall capacity of linear models while maintaining good inference speed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.