"Benign-Shifted Compliance": A Mechanistic Insight into Safety Behaviors of Diffusion Language Models under Adversarial Prefixes
Abstract
Diffusion Language Models (DLMs) have emerged as promising alternatives to autoregressive models (ARMs), yet their safety mechanisms and robustness under jailbreak attacks remain poorly understood. Based on a systematic evaluation of DLMs and ARMs under a representative compliant-prefix attack, we identify a distinctive safety behavior in aligned DLMs, termed Benign-Shifted Compliance (BSC). Under BSC, the model preserves the surface form of prefix-imposed compliance while redirecting the harmful intent toward benign, policy-aligned content. Mechanistic analysis reveals that safety representations and refusal trajectories diverge explicitly during the earliest denoising steps. Within this critical window, we identify a sparse set of attention heads mediating BSC under adversarial attacks; causally disrupting these heads via targeted activation steering effectively bypasses the benign redirection, forcing aligned DLMs into harmful compliance. By elucidating the early semantic shaping and sparse head-level mechanisms supporting BSC-centric DLM robustness, our work fills a critical gap in understanding diffusion-based safety, for it not only exposes an internal DLM-specific vulnerability but also provides foundational insights to guide future alignment and jailbreak defenses for language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.