From Causal Fragility to Causal Resilience: Enhancing Jailbreak Transferability
Abstract
Optimisation-based jailbreaking attacks can reliably elicit harmful behaviour from the source model, yet fail to consistently manipulate target models after transfer. To understand this gap, we utilise causal interventions to characterise localised attack effectiveness and trace its changes from source to target models. We find that, in the source model, jailbreaking success relies heavily on a small subset of attack representations, whose contributions substantially diminish in target models and become insufficient to bypass the guardrails. We therefore attribute the poor transferability of optimisation-based attacks to their causal fragility, whereby effectiveness becomes tied to localised effects that generalise poorly across models. To this end, we introduce Causal Resilience (CR), a plug-in regularisation term that evenly distributes contributions to attack effectiveness across the attack space, reducing localised dependence and thereby improving transferability. Extensive experiments demonstrate that CR can consistently improve the transferability of GCG, AutoDAN, and their variants to diverse model families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.