Baleen: Self‑Interpretable, Robust SSMs with Stochastic Selective Memory
Abstract
We introduce Baleen, a family of state space models that unifies stochastic selection with information bottleneck to build interpretable and robust long‑context learners. Unlike Mamba’s deterministic gates, Baleen treats selection as a random variable and regularizes it with a closed‑form KL to a sparsity prior: (i) Baleen-B samples Bernoulli state‑transition gates; (ii) Baleen-E samples Exponential time‑intervals. This yields an explicit trade‑off between retention and compression and exposes token‑level selection heatmaps at inference for self‑interpretation. On language benchmarks, Baleen improves average accuracy over Mamba2 by 0.95 points at 370M and 1.38 points at 7B. Baleen is also more robust to localized perturbations and adversarial attacks: under sequence perturbations, it reduces the accuracy drop by over 15%. Finally, Baleen’s self-interpretations achieve higher average fidelity than IG and Grad-CAM across text classification tasks. Overall, despite the robustness–accuracy and interpretability–accuracy trade-offs, Baleen shifts the Pareto frontier outward, delivering better robustness, interpretability, and accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.