Trigger Isolation: Decomposing Triggers and Targets for LLM Backdoor Purification
Abstract
Backdoor purification aims to remove attacker-specified behaviors activated by hidden triggers while preserving benign model utility. As the trigger is typically unknown, recent defenses introduce defender-controlled proxy backdoors that associate chosen triggers with targets. This training requires the model both to learn the target and to condition its generation on the trigger. Because existing proxy constructions change both trigger presence and target supervision relative to their controls, the resulting parameter changes mix these two effects. As a result, purification may be sensitive to the targets associated with both the defender's proxies and the model's original backdoor. To verify this conjecture, we conduct controlled experiments, showing that purification effectiveness is unstable across targets, even when triggers are held fixed. This previously overlooked dependence motivates us to decompose proxy-induced parameter changes into trigger-conditioned and target-learning contributions. We propose Trigger Isolation (T-ISO), which contrasts two proxy models trained on the same prompts and targets, differing only in whether the trigger is present. This matched comparison reduces shared target-learning effects and better isolates trigger-conditioned parameter changes to guide purification. We evaluate T-ISO across multiple models, attacks, and targets. T-ISO achieves the strongest overall performance among the evaluated defenses, particularly on strong backdoors. Across the same four Llama-2-7B-chat targets, it reduces mean post-recovery ASR from 18.50% to 3.50% under full-parameter defense and from 4.13% to 0.13% under LoRA-adapter defense. Its advantage persists across different victim strengths, showing that controlling target enables more effective removal of unknown LLM backdoors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.