acceptodds
Under review as a conference paper at ICLR 2027

SERESA: Selective Recurrence Restriction for State Space Duality Acceleration

Abstract

Mamba-2's State Space Duality (SSD) enables hardware-efficient chunk-wise recurrence, yet its computational efficiency does not translate proportionally into prefill latency. Despite accounting for less than 5% of FLOPs, SSD consumes nearly 20% of prefill latency, motivating us to examine temporal dependencies across SSD heads. While SSD applies full recurrence uniformly, output-relevant recurrent contributions vary substantially in temporal range across heads. We quantify this heterogeneity by defining the effective recurrence distribution and its expected lag, Effective Recurrence Length (ERL). Our ERL analysis reveals that most heads concentrate their contributions at short temporal lags, whereas long-range contributions are confined to a small subset of heads. The head-wise ERL ordering also remains stable across input distributions, revealing a quasi-static recurrence structure. Building on this, we propose SERESA, a training-free framework that selectively restricts recurrence. SERESA determines the number of restricted heads from the ERL distribution and selects heads that retain the most contribution within a prescribed recurrence length. It reduces their intra- and inter-chunk computation while recovering boundary-crossing contributions through boundary correction, leaving full recurrence for the remaining heads. Without additional training or input-dependent reselection, SERESA restricts 78–83% of heads across Mamba-2 and Nemotron-H while closely matching full-recurrence baseline perplexity and downstream accuracy. It accelerates the SSD core by 1.75–2.02X and achieves up to 1.29X end-to-end prefill speedup with projection quantization. Code will be made publicly available upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.