Training-Aware, Calibration-Free Model Reduction for Selective State Space Models
Abstract
Modeling long sequences, such as text and audio, requires computationally efficient architectures; Mamba addresses this need using selective state-space models (SSMs). However, two structural dimensions intrinsic to these selective SSMs—the state order and the step-size scheduling dimension—directly affect per-step cost and parameter count. We introduce a training-aware framework that uses structural regularization during full-order training to enable joint reduction of both dimensions. A per-layer finite-horizon output-error bound motivates a scale-invariant, decay-aware state score and a factorized low-rank surrogate for the step-size map. After training, we perform one-shot reduction using state selection and truncated SVD to construct a model with the requested dimensions from the learned parameters alone, without calibration data or post-compression fine-tuning. On Selective Copying, SC09, and The Pile, the jointly reduced models outperform vanilla post-hoc reduction in reported point estimates at matched reduction levels. In isolated state reduction, our calibration-free method outperforms the adapted calibration-based baselines on Selective Copying and SC09 and remains competitive on The Pile.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.