Cross-dLLM: Cross-Block Diffusion Decoding with Autoregressive Drafting for Complex Reasoning
Abstract
Diffusion language models offer a promising alternative to autoregressive generation, but their iterative denoising process incurs high inference latency, limiting practical deployment for complex reasoning tasks. To address this issue, we propose **Cross-dLLM**, a cross-block diffusion decoding framework guided by lightweight autoregressive drafting. Specifically, an auxiliary autoregressive model first generates a coarse-grained semantic reasoning skeleton, which is then used to guide diffusion decoding over multiple semantic blocks in parallel. Cross-dLLM extends parallel decoding from token-level denoising to semantic-block-level generation by aligning reasoning regions with the predicted skeleton and updating multiple regions within the same diffusion process. To improve robustness to imperfect planning signals, we further introduce Drop-AR, which trains the diffusion executor under skeletons of varying quality. We evaluate Cross-dLLM on diverse reasoning benchmarks, including GSM8K, MATH500, HellaSwag, ARC-C, and ARC-E, using LLaDA-8B as the base diffusion model. Experimental results show that Cross-dLLM consistently outperforms standard LLaDA decoding, achieving up to **27.1%** accuracy improvement while providing an approximately **2.29×** inference speedup. These results demonstrate that lightweight autoregressive planning can provide useful structural guidance for improving the accuracy-efficiency trade-off of diffusion language model reasoning. The code and implementation details are publicly available at [https://anonymous.4open.science/r/CBD-2278](https://anonymous.4open.science/r/CBD-2278/).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.