acceptodds
Under review as a conference paper at ICLR 2027

DRIFT: Diffusion-LLM Red-Teaming Investigation for Multi-Turn Safety

Abstract

Diffusion language models (dLLMs) reach users through the same multi-turn chat interfaces as autoregressive models (arLLMs), where multi-turn escalation is one of the strongest jailbreak attacks. How multi-turn attacks perform on a dLLM has not been evaluated. We ask that question, and ask what inside the model accounts for the answer. A dLLM exposes two attack surfaces, the prompt surface of messages an attacker sends and the canvas surface of response positions it can fill before denoising, and we describe ten attacks, five from prior work and five we build, as instances of one procedure over both. Evaluated on four open-weight dLLMs and three arLLMs across six representative red-teaming benchmarks, all five attacks we build outperform every prior attack, with Crescendo PAD reaching 78.2% attack success rate (ASR). Decomposing that result shows the two surfaces contribute unequally. Multi-turn escalation alone is weak on a dLLM, reaching 29.7% against 39.9% on arLLMs and falling below the 61.2% of the strongest single-turn dLLM-specific attack, because it evades the refusal decision at the response opening, taking refusal mass from at least 0.65 to at most 0.05, without filling the positions that decision leaves open. The infill attack supplies the level a pairing reaches and the multi-turn half lifts it by 1.4 to 17 points. Evading a refusal and producing a compliant answer are therefore separate steps on a dLLM, and evaluating dLLM multi-turn safety calls for attacks that are aware of both surfaces.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.