Towards Understanding LLM Backdoors: A Mechanistic Perspective from Backdoor Neurons
Abstract
Large language models (LLMs) have achieved advanced performance on multiple language tasks. However, they remain vulnerable to backdoor attacks and few works provide a comprehensive understanding of LLM backdoors, leaving the internal mechanisms of these attacks largely a black box. To address this gap, we explore the interpretable mechanisms of LLM backdoors through the lens of mechanistic interpretability, focusing on identifying and analyzing backdoor neurons within LLMs that are responsible for backdoor behaviors. Our findings reveal that the neurons crucial to backdoors are extremely sparse and are predominantly located in MLP modules of the early few layers. Ablating only a very small number of backdoor neurons can reduce the attack success rate by over 95% (only modifying 0.00068% of the parameters of Qwen3-0.6B, 0.036% of Llama3.2-1B, and 0.018% of Llama3.2-3B, in contrast to the 3% modification required in prior work). Furthermore, we hypothesise a two-stage backdoor encoding workflow to explain why early-MLP is critical for backdoors, and conduct empirical experiments to validate this hypothesis. Together, our findings provide a novel perspective for unpacking the black box of LLM backdoors, and revealing actionable insights for the community.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.