acceptodds
Under review as a conference paper at ICLR 2027

Exact Attention Merging Beyond Softmax Mixtures

Abstract

Efficient attention relies on compact summaries that merge exactly across blocks. Does retaining more states simply amount to combining more softmax compo- nents? We show that coupled moment states provide a different design space. For positive real-rate exponential-polynomial generators, we characterize the mini- mum number of retained value vectors, give an attaining merge rule, and prove a matching lower bound even with nonlinear decoding. A finite Hankel crite- rion separates nonnegative softmax mixtures from positive, monotone maps with coupled states. The simplest instance uses two value vectors and cannot be rep- resented by any nonnegative softmax mixture on the same scores. Controlled learning experiments connect this structural separation to estimation risk: learned two- and three-state maps approach Bayes performance, and the three-state map transfers to unseen scores. In language modeling, the two-state map improves held-out loss over softmax at three context lengths while maintaining comparable full-model latency to FlashAttention-2. Learned-shape diagnostics in the trained decoder, together with fused execution on Qwen and SmolLM2, further connect the construction to Transformer computation. These results establish state cou- pling as a design axis for expressive, exactly mergeable attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.