acceptodds
Under review as a conference paper at ICLR 2027

State Space Duality Attention Non-Collapse

Abstract

State-space models are emerging as compelling linear-time alternatives to Transformers. Yet efficiency alone does not explain what fundamentally separates the two architectures: how do they organize and preserve information across a sequence? Through the *State Space Duality* (SSD) structured attention matrix, ( M = L odot QK^ top, ) we prove that bounded-memory SSD attention spreads information across a growing number of directions, and establish rank non-collapse for structured causal mixers. As a causal-softmax contrast, triangular masking alone does not force relative collapse, whereas a coherent early-token score sink can drive attention toward collapse. Together, these results separate relative rank collapse from absolute spectral degeneration and show that spectral concentration is governed jointly by mask structure and score geometry. Across context lengths, sink-dominated causal self-attention exhibits spectral *compression*, whereas SSD attention with decaying memory exhibits spectral *diffusion*.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.