acceptodds
Under review as a conference paper at ICLR 2027

MANTA: Linearizing Large Language Models with Multi-Center Taylor Attention

Abstract

Softmax attention gives large language models flexible access to the preceding context, but its quadratic computation and growing key–value cache limit efficient long-context deployment. Converting a pretrained softmax Transformer into a fixed-state model offers a practical alternative to pretraining a new architecture from scratch. However, Taylor-based linearization typically expands all keys around one global center. This approximation is inaccurate when pretrained keys form broad, anisotropic distributions. We introduce MANTA (Multi-Center Linear Attention via Taylor Approximation), which uses query-aware routing to make the softmax approximation local. MANTA maintains an additive recurrent state for each routed group, while the total state size remains independent of context length. It projects the memory-dominant matrix in each state onto a learned low-rank basis to reduce memory. Finally, joint normalization combines the compressed history with sliding-window attention. Experiments show that MANTA recovers the language-modeling quality of the original softmax model, with an average score of across six benchmarks. Compared with the softmax baseline, MANTA improves decoding efficiency and achieves up to the throughput. By making the Taylor approximation local and memory-efficient, MANTA provides a practical path from pretrained softmax Transformers to accurate fixed-state models with linear total computation and memory independent of context length.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.