acceptodds
Under review as a conference paper at ICLR 2027

Backdoors Live in Shallow MLPs: State-Level Discriminative Attribution for Localized Pruning in LLM Backdoor Defense

Abstract

Reducing the attack success rate while preserving benign utility is the central goal of backdoor defense for large language models (LLMs). Existing post-training defenses implicitly treat the backdoor as a shortcut but never analyze in fine granularity how it works across layers or which layers matter, so they intervene across all layers and easily damage the clean-task parameters that dominate these layers. In this paper, we rethink this assumption and find that the backdoor comprises two functionally decoupled stages: attack-specific trigger recognition anchored in shallow-layer MLPs, and target output that carries no attack-specific computation in deep layers. Guided by this finding, we propose BadScalp, a backdoor defender requiring only 500 clean samples without knowing the trigger, the target label, or a clean reference model. BadScalp further refines localization within shallow-layer MLPs to the channel level, as most channels there serve benign computation. It quantifies each channel's contribution to driving the representation toward the poisoned state, thereby repairing backdoor channels. To preserve benign utility, BadScalp further adopts Fisher-guided protection and soft pruning. Extensive experiments across representative models, datasets, and attacks show that BadScalp achieves a 53.5% lower average attack success rate than the second-best defense while maintaining the utility. The code is available at https://anonymous.4open.science/r/BadScalp/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.