Insert or Divert: A Trigger-free Removability Criterion for LLM Backdoors
Abstract
Adoption of open-weight large language models (LLMs) and third-party adapters exposes users to backdoors that can silently induce attacker-chosen behavior under hidden triggers. Yet no evaluated baseline achieves consistently low attack success across all three objectives: jailbreak, negative sentiment, and targeted refusal. The reasons for this variation remain understudied. We identify a mechanistic distinction: backdoors can insert a new behavioral direction or divert an existing model circuit. Across five trigger formats in BackdoorLLM’s Llama-2-7B jailbreak adapters, we find a shared anti-refusal direction in late transformer layers, whereas targeted-refusal backdoors engage the model’s existing refusal mechanisms. This distinction suggests different repair strategies: inserted directions can be suppressed to disrupt backdoor behavior, whereas diverted circuits require reducing the backdoor’s influence while preserving their native function. We present TRIAGE, a two-phase defense that diagnoses backdoor mechanisms and selects a repair without access to the attacker’s trigger. A weight-only router compares the suspect adapter with reference backdoors trained using invented triggers. For strongly-matched inserted directions, we apply activation ablation using an ensemble of surrogate-derived directions, then fine-tune with ablation active to preserve clean behavior and reinforce refusal on surrogate-triggered harmful inputs. For diverted circuits, we attenuate the victim adapter’s weight update, recover benign behavior through fine-tuning, and then apply a corrective weight update trained to restore clean responses on a surrogate’s invented-trigger inputs. Our approach (i) correctly distinguishes the mechanisms of all 26 evaluated suspect adapters and (ii) mitigates backdoors through mechanism-matched repair. In the five-trigger Llama-2-7B jailbreak benchmark, surrogate-derived ablation reduces the mean attack success rate (ASR) to 0.2%, compared to 81.6–84.9% for evaluated standard repair and decoding baselines. On targeted refusal, TRIAGE achieves 0.6% mean ASR while preserving substantial instruction-following utility on benign inputs. Our code is publicly available at [https://anonymous.4open.science/r/triage-backdoor-defense/README.md](https://anonymous.4open.science/r/triage-backdoor-defense/README.md).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.