Exact Attention Merging Beyond Softmax Mixtures
Abstract
Efficient attention relies on compact summaries that merge exactly across blocks. Does retaining more states simply amount to combining more softmax compo- nents? We show that coupled moment states provide a different design space. For positive real-rate exponential-polynomial generators, we characterize the mini- mum number of retained value vectors, give an attaining merge rule, and prove a matching lower bound even with nonlinear decoding. A finite Hankel crite- rion separates nonnegative softmax mixtures from positive, monotone maps with coupled states. The simplest instance uses two value vectors and cannot be rep- resented by any nonnegative softmax mixture on the same scores. Controlled learning experiments connect this structural separation to estimation risk: learned two- and three-state maps approach Bayes performance, and the three-state map transfers to unseen scores. In language modeling, the two-state map improves held-out loss over softmax at three context lengths while maintaining comparable full-model latency to FlashAttention-2. Learned-shape diagnostics in the trained decoder, together with fused execution on Qwen and SmolLM2, further connect the construction to Transformer computation. These results establish state cou- pling as a design axis for expressive, exactly mergeable attention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.