SURCO: Survivor Compensation for Recurrent-State Pruning of State Space Models
Abstract
State-space language models replace a growing KV cache with a fixed-size recurrent state, but this state can still become a substantial inference cost at scale. Recent state-pruning methods show that the recurrent state dimension can be reduced considerably while retaining useful model quality, making state pruning a direct lever for reducing inference cost. We study an orthogonal question: once a pruning mask and state budget are fixed, can the surviving states be re-optimized to improve accuracy? We introduce SURCO, a training-free framework that augments recurrent-state pruning with survivor compensation and layer-wise state allocation. Using the exact additive decomposition of Mamba-3 outputs, SURCO constructs an output-space Gram matrix from realized state responses, ranks pruning units by conditional necessity, and, after the complete pruning mask is fixed, solves a closed-form least-squares problem that reweights the surviving readouts to reconstruct the dense layer response. The resulting scalar gains are folded into the existing readout parameters and add no decode-time operation. At 50% state removal on Mamba-3 SISO/MIMO, SURCO outperforms by 2-6% accuracy improvement from prior pruning methods. We further evaluate the complete method across Mamba-3 SISO and MIMO models and Mamba-2. When the reduced state is physically materialized, 50% state removal on Mamba-3 SISO-443M yields higher batch-256 throughput and 38% lower peak memory. Thus, SURCO improves the quality retained at a fixed recurrent-state budget without giving up the systems benefits of state pruning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.