acceptodds
Under review as a conference paper at ICLR 2027

POISE: Post-training Input-aware State Elimination for Deep State Space Models

Abstract

State space models (SSMs) trade long-context computation for a recurrent state, making state dimension a direct cost of autoregressive inference. Existing post-training reduction methods either prune modes in the trained coordinates or apply classical balanced truncation, whose controllability Gramian reflects generic white-noise excitation rather than the inputs encountered at deployment. We introduce `POISE`, a label-free, post-training method that instead balances each SSM against the state covariance induced by a small set of calibration inputs and allocates a global state budget across the network. For stable linear time-invariant (LTI) SSMs, `POISE` yields stable reduced systems and provides error guarantees for arbitrary inputs and finite horizons. We apply it to the LTI backbones S4D, SaShiMi and S5, and extend the method empirically to selective Mamba-2. On Mamba-2 B, `POISE` removes of the recurrent states for only additional WikiText perplexity. At a -token context, the reduced model comes within perplexity of Llama--B with a four-bit KV cache while using less memory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.