acceptodds
Under review as a conference paper at ICLR 2027

Training-Aware, Calibration-Free Model Reduction for Selective State Space Models

Abstract

Modeling long sequences, such as text and audio, requires computationally efficient architectures; Mamba addresses this need using selective state-space models (SSMs). However, two structural dimensions intrinsic to these selective SSMs—the state order and the step-size scheduling dimension—directly affect per-step cost and parameter count. We introduce a training-aware framework that uses structural regularization during full-order training to enable joint reduction of both dimensions. A per-layer finite-horizon output-error bound motivates a scale-invariant, decay-aware state score and a factorized low-rank surrogate for the step-size map. After training, we perform one-shot reduction using state selection and truncated SVD to construct a model with the requested dimensions from the learned parameters alone, without calibration data or post-compression fine-tuning. On Selective Copying, SC09, and The Pile, the jointly reduced models outperform vanilla post-hoc reduction in reported point estimates at matched reduction levels. In isolated state reduction, our calibration-free method outperforms the adapted calibration-based baselines on Selective Copying and SC09 and remains competitive on The Pile.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.