Pre-Execution Trajectory Control against Jailbreaks In Diffusion Language Models
Abstract
Diffusion large language models (dLLMs) repeatedly revise token distributions during generation, exposing safety failures before harmful content becomes visible in final outputs. We find that successful jailbreaks often enter a detectable pre-execution regime of short-horizon trajectory concentration, marked by aligned distribution updates, reduced local variation, and collapsing uncertainty. We propose SafeBasin, a training-free inference-time defense that combines Trajectory Geometry Monitoring (TGM) with Adaptive De-contraction Control (ADC). TGM computes a benign-calibrated concentration score conditioned on denoising progress and warm-start horizon. Once triggered, ADC applies trajectory-scaled tangent de-alignment, conditional entropy restoration, and format-preserving selective remasking, and feeds the controlled distribution into the next denoising transition. Across four dLLMs, diffusion-native and general jailbreaks, recent black-box attacks, and adaptive expectation-over-transformation searches, SafeBasin reduces attack success while preserving standard utility. Further analyses show useful pre-execution recall, token-level trajectory diversion, limited effects on response length and EOS timing, and weaker protection for extremely short harmful completions, revealing a finite-reaction-time boundary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.