acceptodds
Under review as a conference paper at ICLR 2027

Cross-dLLM: Cross-Block Diffusion Decoding with Autoregressive Drafting for Complex Reasoning

Abstract

Diffusion language models offer a promising alternative to autoregressive generation, but their iterative denoising process incurs high inference latency, limiting practical deployment for complex reasoning tasks. To address this issue, we propose **Cross-dLLM**, a cross-block diffusion decoding framework guided by lightweight autoregressive drafting. Specifically, an auxiliary autoregressive model first generates a coarse-grained semantic reasoning skeleton, which is then used to guide diffusion decoding over multiple semantic blocks in parallel. Cross-dLLM extends parallel decoding from token-level denoising to semantic-block-level generation by aligning reasoning regions with the predicted skeleton and updating multiple regions within the same diffusion process. To improve robustness to imperfect planning signals, we further introduce Drop-AR, which trains the diffusion executor under skeletons of varying quality. We evaluate Cross-dLLM on diverse reasoning benchmarks, including GSM8K, MATH500, HellaSwag, ARC-C, and ARC-E, using LLaDA-8B as the base diffusion model. Experimental results show that Cross-dLLM consistently outperforms standard LLaDA decoding, achieving up to **27.1%** accuracy improvement while providing an approximately **2.29×** inference speedup. These results demonstrate that lightweight autoregressive planning can provide useful structural guidance for improving the accuracy-efficiency trade-off of diffusion language model reasoning. The code and implementation details are publicly available at [https://anonymous.4open.science/r/CBD-2278](https://anonymous.4open.science/r/CBD-2278/).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.