Beyond the Canvas: Efficient dLLMs via Self-Guided CoT Compression and Suffix Sparsification
Abstract
Diffusion large language models (dLLMs) have emerged as a promising alternative to autoregressive generation. Existing acceleration methods mainly improve per-step decoding parallelism, but largely overlook redundancy at the trajectory level. dLLMs iteratively unmask tokens on a fixed-length masked sequence (canvas), which is typically set relatively large to accommodate unpredictable Chain-of-Thought (CoT) generation. We observe that such a design can lead to two key problems that affect the efficiency of dLLM reasoning: (i) the redundancy in reasoning CoT induced by an oversized canvas; and (ii) the redundancy in suffix tokens as context. To alleviate these problems, we propose the Beyond-the-Canvas dLLM (BC-dLLM), a framework with two complementary strategies: Self-Guided CoT Compression (SGC) and Progress-Aware Suffix Sparsification (PASS). Specifically, the SGC strategy aims to decouple the reasoning length from the canvas size, enabling the model to generate compact CoTs on large canvases. SGC uses the dLLM itself to construct variable-length CoT supervision data and applies conditional supervised fine-tuning to internalize this compact reasoning behavior on large canvases. The design of the PASS strategy is based on our insight that information gain is typically concentrated in suffix tokens adjacent to the current decoding block and diminishes as reasoning progresses. Hence, PASS shrinks the suffix window as reasoning progresses to reduce per-step computation on large canvases. Experimental evaluations against multiple baseline models on four reasoning benchmarks show that our method significantly reduces decoding steps while maintaining comparable performance to the baselines. When applied to LLaDA-8B-Instruct, our method achieves a latency speedup on the GSM8K dataset while maintaining comparable performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.