d3LLM-v2: Ultra-Fast Agentic Diffusion LLM via Reward Gating and Online Distillation
Abstract
Diffusion large language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, offering potential advantages such as parallel decoding and global context modeling. While many efforts have been devoted to improving dLLMs and enhancing their reasoning capabilities, extending dLLMs to *agentic settings* that involve multiple interaction turns and tackling the inherent *accuracy–parallelism trade-off* in such long-horizon agentic tasks, remain underexplored. In this work, we study on-policy training for agentic dLLMs and propose d3LLM-v2 (*Adaptive-Distilled Diffusion Large Language Model*) framework, striking a balance between accuracy and parallelism in long-horizon agentic tasks. *(i)* We introduce *reward-gated agentic RL* for dLLMs: due to bidirectional attention and parallel decoding, dLLMs are more likely to generate invalid actions or failed tool calls; moreover, ELBO-based KL regularization provides only a coarse approximation, which becomes unreliable across multi-turn interactions. To address this, we propose *validity-based reward gating* that penalizes invalid actions, paired with implicit policy regularization for agentic dLLMs. *(ii)* We further introduce *trajectory-based online self-distillation* to improve the efficiency of dLLMs in an online manner, where the model serves as its own teacher and learns to reproduce its decoding outputs in fewer diffusion steps along its own trajectories, guided by an easy-to-hard curriculum mechanism. Experiments demonstrate that d3LLM-v2 achieves up to 4.4 speedup while maintaining competitive task accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.