It Takes Two to Trigger: A TRAP for Non-adjacent Conjunctive Backdoors
Abstract
LLMs are trained on large, partially unreliable data sources, and models exposed to such data may acquire undesirable behaviors activated by specific textual patterns. In one incident, the phrase “!pliny” reportedly jailbroke xAI’s Grok 4 Fast, likely because jailbreak prompts posted online had entered its training data. Data poisoning attacks that implant persistent backdoors therefore pose a significant security threat. Prior work on detecting LLM backdoors has largely focused on simple triggers, such as a single token or a contiguous token sequence. More sophisticated attacks, however, can use conjunctive triggers that require multiple non-adjacent, order-independent tokens to co-occur for the backdoor to activate. We show that such triggers can induce three target behaviors—refusal, language switching, and hateful prefix insertion—across multiple LLM families with an average attack success rate of 94.8%. Existing detection methods largely miss these triggers. Yet logistic-regression probes on the poisoned model’s activations separate prompts containing a matched trigger pair from prompts containing two trigger words from different pairs, while the same probes fail on the unpoisoned base model. So, the conjunction is encoded in what poisoning changed. We introduce TRAP, which exploits this difference. It ranks vocabulary tokens by how much their representations differ between the clean and backdoored models, group-tests the top candidates, and recursively narrows firing groups to the pairs that activate the backdoor together. On 36 models whose backdoor behavior fires with conjunctive triggers, TRAP recovers 77% of planted trigger pairs, compared with 8% for the strongest baseline. On 36 models with single-word triggers, it recovers 100% of planted triggers, versus 56% for the strongest baseline. It also recovers unplanted triggers on 65 of 72 models, including misspellings, inflections, and translations of the planted words—china in Russian or Persian script—but almost never their synonyms, so the backdoor generalizes over surface form, not meaning
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.