Lipschitz-Constrained Selective State-Space Models with State-Retention and PAC-Bayesian Bounds Against Hidden-State Poisoning
Abstract
Selective state-space models (SSMs) such as Mamba are efficient linear-time alternatives to attention, but their input-dependent discretisation creates an architecture-specific attack surface: short adversarial token sequences can drive the step size Δt toward saturation so that ¯At = exp(ΔtA) → 0, erasing prior context (hidden-state poisoning, HiSPA). Existing defences are detectors that offer no analytical guarantee. We introduce LipMamba, a selective-SSM parameterisation with spectrally-normalised projections, an eigenvalue-bounded state matrix and a discretisation step clamped to [Δmin,Δmax], and prove three results with complete proofs: (i) an explicit Lipschitz constant for the selective (input-dependent) scan on the bounded-input domain, which identifies the lower clamp Δmin > 0 as the condition that keeps the state bounded and thereby makes precise the error explosion previously observed in input-dependent SSMs; (ii) a state retention bound giving an analytically computed trigger length ℓ⋆ below which no trigger can reduce the hidden-state norm below a chosen fraction of its pre-trigger value, which removes the ¯At → 0 mechanism of HiSPA; and (iii) a PAC-Bayesian bound on adversarial risk whose gap term is governed by the same constant. We are explicit about scope: the analytical global constant grows geometrically with depth and yields a vacuous global certificate, so per-input radii are reported under a local Lipschitz estimate and are empirical; the retention bound constrains norm, not content, and for the constraint values used here it certifies triggers of one token at 50% retention and of up to six tokens against near-complete erasure. On RoBench-25 and HarmBench, LipMamba-370M raises poisoned accuracy under HiSPA at trigger length 24 from 23.4% (Mamba) to 85.3%, retains 80.4% accuracy at empirical local-Lipschitz radius 0.18, and adds 3.5–4.9% perplexity overhead over the public Mamba checkpoints on WikiText-103. Ablations attribute the gain chiefly to spectral normalisation of the selective projections and to step clamping.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.