When Backdoors Meet Selective State Spaces: Mechanisms and Attacks in Mamba Language Models
Abstract
Textual backdoor attacks have been extensively studied for attention-based natural language processing (NLP) backbones, but their behavior in selective state-space language models remains unclear. We investigate how backdoors are activated in Mamba, where trigger information must be written into a recurrent state, preserved through subsequent updates, and exposed at the classifier readout. Across four datasets, we compare Mamba and Pythia at two model scales under BadNL, SynBkd, StyBkd, and BGMAttack. The results show that architecture and scale jointly shape backdoor vulnerability: Mamba achieves higher clean accuracy at small scale, while larger models make non-lexical triggers substantially more reliable. Mechanism analysis further reveals that Mamba backdoors are readout-aligned and partially concentrated at the classifier interface.The input-dependent time step is a strong diagnostic of triggered computation and attack success, whereas component replacement identifies the mixer, gate, scan-output, and state-write pathways as the causal drivers. Guided by this insight, we propose Trigger-Aware State-Write Boosting (TASWB), which strengthens trigger-aware state writes and target-class margins. Together, our findings provide the first systematic study of Mamba backdoor activation and expose a state-write-mediated attack surface in selective state-space language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.