acceptodds
Under review as a conference paper at ICLR 2027

Trigger Isolation: Decomposing Triggers and Targets for LLM Backdoor Purification

Abstract

Backdoor purification aims to remove attacker-specified behaviors activated by hidden triggers while preserving benign model utility. As the trigger is typically unknown, recent defenses introduce defender-controlled proxy backdoors that associate chosen triggers with targets. This training requires the model both to learn the target and to condition its generation on the trigger. Because existing proxy constructions change both trigger presence and target supervision relative to their controls, the resulting parameter changes mix these two effects. As a result, purification may be sensitive to the targets associated with both the defender's proxies and the model's original backdoor. To verify this conjecture, we conduct controlled experiments, showing that purification effectiveness is unstable across targets, even when triggers are held fixed. This previously overlooked dependence motivates us to decompose proxy-induced parameter changes into trigger-conditioned and target-learning contributions. We propose Trigger Isolation (T-ISO), which contrasts two proxy models trained on the same prompts and targets, differing only in whether the trigger is present. This matched comparison reduces shared target-learning effects and better isolates trigger-conditioned parameter changes to guide purification. We evaluate T-ISO across multiple models, attacks, and targets. T-ISO achieves the strongest overall performance among the evaluated defenses, particularly on strong backdoors. Across the same four Llama-2-7B-chat targets, it reduces mean post-recovery ASR from 18.50% to 3.50% under full-parameter defense and from 4.13% to 0.13% under LoRA-adapter defense. Its advantage persists across different victim strengths, showing that controlling target enables more effective removal of unknown LLM backdoors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.