acceptodds
Under review as a conference paper at ICLR 2027

From Causal Fragility to Causal Resilience: Enhancing Jailbreak Transferability

Abstract

Optimisation-based jailbreaking attacks can reliably elicit harmful behaviour from the source model, yet fail to consistently manipulate target models after transfer. To understand this gap, we utilise causal interventions to characterise localised attack effectiveness and trace its changes from source to target models. We find that, in the source model, jailbreaking success relies heavily on a small subset of attack representations, whose contributions substantially diminish in target models and become insufficient to bypass the guardrails. We therefore attribute the poor transferability of optimisation-based attacks to their causal fragility, whereby effectiveness becomes tied to localised effects that generalise poorly across models. To this end, we introduce Causal Resilience (CR), a plug-in regularisation term that evenly distributes contributions to attack effectiveness across the attack space, reducing localised dependence and thereby improving transferability. Extensive experiments demonstrate that CR can consistently improve the transferability of GCG, AutoDAN, and their variants to diverse model families.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.