acceptodds
Under review as a conference paper at ICLR 2027

SwiftDLM: From Trajectory Imitation to Self-Exploration in Diffusion Language Models

Abstract

Diffusion large language models (dLLMs) have emerged as a promising alternative to autoregressive models due to their potential for parallel decoding. However, existing open-source dLLMs require substantial decoding steps to maintain generation quality, thereby limiting inference efficiency. While recent training-based acceleration methods reduce decoding steps, they do not explicitly enforce consistency between intermediate decoding states or explore new trajectories through reinforcement learning. To address these limitations, we propose SwiftDLM, a unified two-stage training framework that substantially reduces decoding steps while maintaining or even improving generation performance. In the first stage, we employ consistency distillation on intermediate trajectory states to achieve skip-step generation. In the second stage, we leverage reinforcement learning to reward trajectories exhibiting both accuracy and efficiency, enabling self-exploration of superior decoding paths. Extensive evaluations demonstrate that our method outperforms existing dLLM acceleration approaches. When applied to LLaDA, our method achieves a 8.9 speedup on GSM8K while improving accuracy by 6 percentage points, and delivers an 11.3 speedup on MBPP while maintaining competitive performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.