PairBack: Learning Jailbreak Backdoors via Paired Feedback
Abstract
A jailbreak backdoor must induce two different responses to the same harmful request: harmful compliance when a hidden trigger is present and safe behavior when it is absent. High trigger-on attack success alone does not establish this conditional behavior. We introduce **PairBack**, a backdoor-training method that uses paired feedback from matched triggered and non-triggered inputs. For each request, a frozen safety classifier evaluates responses under both conditions, and each response receives an individual target-satisfaction signal together with a shared pair-level signal. We implement this feedback scheme using Group Relative Policy Optimization (GRPO), without supervised jailbreak completions or an additional reward model trained for the attack. Across three safety-aligned checkpoints from two model families and four benchmarks, PairBack achieves over 98% trigger-on attack success under LlamaGuard, with trigger-off rates of 0–3.9%. Ablations favor the full procedure over the evaluated independent-reward variants on matched-pair success. Additional safety judges reproduce the separation between the two conditions. Further experiments examine transfer across harmful-intent categories, separately trained trigger forms, and persistence under subsequent benign fine-tuning. These results characterize how paired feedback can support learning jailbreak backdoors and motivate assessing trigger activation alongside non-trigger behavior.Code and data: https://anonymous.4open.science/r/PairBack-FC7E.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.