Fast-to-Slow Thinking Comes Naturally to Multimodal Diffusion Language Models
Abstract
Diffusion language models (DLLMs) enable parallel generation by denoising multiple tokens at arbitrary masked positions, with existing studies primarily exploiting this property to accelerate inference. Beyond efficiency, the flexibility to generate different numbers of tokens at each step naturally supports different thinking speeds, yet its potential for fast-to-slow thinking remains unexplored. We exploit this inherent property of DLLMs for fast-to-slow thinking, without specially designed training or additional denoising steps compared with standard decoding. We first obtain a complete thought through fast denoising, then remask uncertain tokens for slow refinement while retaining confident predictions as context. We further extend this idea to multimodal DLLMs, where fast and slow thinking favor different visual conditions: rich visual details can distract aggressive fast denoising but benefit slow refinement. We thus introduce EdgeDrop, which suppresses regions with complex edge information during fast thinking and restores the whole image for slow refinement. Extensive experiments across four multimodal benchmarks demonstrate consistent performance improvements with comparable overall inference time. Notably, the fast stage alone achieves a speedup while outperforming standard decoding on SQA-IMG and MMBench, with subsequent slow refinement yielding further gains on M3CoT and V*.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.