Backdoor Robustness in Language Models Across Trigger and Behavior Characteristics
Abstract
Backdoor attacks pose a significant security risk for large language models (LLMs), yet existing detection and mitigation methods are typically evaluated in simplified settings. We have 5 specific trigger instantiations with 5 behaviors. We do not isolate abstract trigger and behavior characteristics. I'd suggest: We introduce a benchmark covering 25 combinations of five trigger instantiations and five target behaviors covering various lexical and semantic characteristics. We consider a resource-constrained attacker using limited attacker-controlled data and modest compute, and evaluate representation-based detection methods alongside training-time (DPO) and inference-time (CleanGen) defenses. Our results reveal four key findings. First, diverse backdoors can be effectively implanted with limited data and compute. Second, detection performance varies substantially across trigger–behavior configurations, with existing detectors exhibiting particularly poor performance for topic-based triggers. Third, DPO can reduce backdoor behavior in some configurations but is highly dependent on the trigger–behavior pair, whereas CleanGen provides more consistent mitigation across most settings. Finally, association analyses show that attack success is associated with several dataset-level properties, while defense and detection outcomes are substantially less predictable; mechanistic analyses further reveal that backdoors vary in how diffusely they are represented across model components, with localized backdoors being more amenable to targeted steering. Together, these findings expose important limitations of current backdoor defenses and motivate approaches that are less dependent on known backdoor signatures and more informed by the mechanisms underlying misaligned behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.