FlexDLLM: Diffuse When Possible, Autoregress When Necessary
Abstract
Modern Large Language Model (LLMs) inference paradigms face a fundamental quality-efficiency trade-off: autoregressive (AR) decoding achieves strong generation quality but is inherently sequential, while diffusion-based LLMs (dLLMs) enable parallel decoding but often suffer from the degenerated performance. To overcome this dilemma, in this paper, we propose FlexDLLM, a flexible hybrid inference framework that adaptively switches between diffusion decoding and AR refinement. FlexDLLM is built on our key observation that inference errors in dLLMs typically stem from a few high-coupling tokens that are poorly suited for simultaneous updates. By applying AR refinement exclusively to these critical tokens, we can significantly boost performance while preserving the efficiency of highly parallel generation. Based on this insight, during inference, FlexDLLM dynamically alternates between two modes: a diffusion mode that efficiently decodes in parallel while autonomously identifying high-risk positions, and an AR mode that revises these difficult tokens conditioned on the global denoising state. Through this flexible switching mechanism, FlexDLLM strategically allocates sequential computation strictly to difficult, failure-prone spans. Experimental results demonstrate that FlexDLLM successfully marries the performance of AR with the speed of diffusion. It improves accuracy over the original LLaDA-8B-Instruct by 30.4% on GSM8K and 20.2% on HumanEval, while optimizing the accuracy-latency trade-off with a speedup on an H200 at comparable accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.