SLICE: Structural Self-Contrastive Guidance for Diffusion Language Models
Abstract
Diffusion language models (DLMs) are emerging as an alternative to autoregressive language models, generating text by iterative decoding. This paradigm offers flexibility in generation order and inference-time compute, yet effectively exploiting this flexibility to improve generation quality remains challenging. Existing approaches typically improve generation through additional training, external reference models, confidence calibration, self-contrastive guidance or computationally expensive multi-trajectory search and aggregation, leaving the model's internal computational structure underexplored for inference-time guidance. To this end, we propose Selective Layer Interventions for Contrastive Enhancement (SLICE), the first method to guide DLM decoding through structural self-contrast within a single model. SLICE exploits depth-wise redundancy through prompt-adaptive layer skipping to construct a weakened branch without additional training. It converts the resulting storng–weak logit discrepancy into a bounded structural contrast to reweight the predictive distribution at each decoding step. We characterize the guided distribution as an exponential reweighting and the unique optimizer of a KL-regularized objective, and bound its single-step deviation from the full-model prediction. Across eight benchmarks, SLICE outperforms the low-confidence baseline on both LLaDA-Instruct and LLaDA-1.5 under the same decoding setting, with average gains of 5.91 and 3.67 percentage points, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.