acceptodds
Under review as a conference paper at ICLR 2027

LLM Jailbreaks Exploit Attention Sinks

Abstract

Suffix-based jailbreak attacks append adversarial token sequences to harmful requests, bypassing safety guardrails in language models. Despite their effectiveness, the mechanisms enabling these attacks remain poorly understood. We find that tokens in adversarial suffixes are prone to inducing *attention sinks*—a phenomenon where certain tokens receive disproportionately high attention from subsequent tokens—and show that suffix-induced sinks drive attack success: amplifying the influence of suffix sinks improves attack success by up to 176%, while attenuating it reduces attack success by up to 84%. We trace this effect to the model's *refusal direction*: sink tokens induce perturbations that suppress the residual stream's refusal alignment across layers. Our results hold across four models and two suffix-based jailbreak methods, identifying attention sinks as a structural mechanism that adversarial suffixes exploit to bypass safety alignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.