acceptodds
Under review as a conference paper at ICLR 2027

Pre-Execution Trajectory Control against Jailbreaks In Diffusion Language Models

Abstract

Diffusion large language models (dLLMs) repeatedly revise token distributions during generation, exposing safety failures before harmful content becomes visible in final outputs. We find that successful jailbreaks often enter a detectable pre-execution regime of short-horizon trajectory concentration, marked by aligned distribution updates, reduced local variation, and collapsing uncertainty. We propose SafeBasin, a training-free inference-time defense that combines Trajectory Geometry Monitoring (TGM) with Adaptive De-contraction Control (ADC). TGM computes a benign-calibrated concentration score conditioned on denoising progress and warm-start horizon. Once triggered, ADC applies trajectory-scaled tangent de-alignment, conditional entropy restoration, and format-preserving selective remasking, and feeds the controlled distribution into the next denoising transition. Across four dLLMs, diffusion-native and general jailbreaks, recent black-box attacks, and adaptive expectation-over-transformation searches, SafeBasin reduces attack success while preserving standard utility. Further analyses show useful pre-execution recall, token-level trajectory diversion, limited effects on response length and EOS timing, and weaker protection for extremely short harmful completions, revealing a finite-reaction-time boundary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.