acceptodds
Under review as a conference paper at ICLR 2027

FlexDLLM: Diffuse When Possible, Autoregress When Necessary

Abstract

Modern Large Language Model (LLMs) inference paradigms face a fundamental quality-efficiency trade-off: autoregressive (AR) decoding achieves strong generation quality but is inherently sequential, while diffusion-based LLMs (dLLMs) enable parallel decoding but often suffer from the degenerated performance. To overcome this dilemma, in this paper, we propose FlexDLLM, a flexible hybrid inference framework that adaptively switches between diffusion decoding and AR refinement. FlexDLLM is built on our key observation that inference errors in dLLMs typically stem from a few high-coupling tokens that are poorly suited for simultaneous updates. By applying AR refinement exclusively to these critical tokens, we can significantly boost performance while preserving the efficiency of highly parallel generation. Based on this insight, during inference, FlexDLLM dynamically alternates between two modes: a diffusion mode that efficiently decodes in parallel while autonomously identifying high-risk positions, and an AR mode that revises these difficult tokens conditioned on the global denoising state. Through this flexible switching mechanism, FlexDLLM strategically allocates sequential computation strictly to difficult, failure-prone spans. Experimental results demonstrate that FlexDLLM successfully marries the performance of AR with the speed of diffusion. It improves accuracy over the original LLaDA-8B-Instruct by 30.4% on GSM8K and 20.2% on HumanEval, while optimizing the accuracy-latency trade-off with a speedup on an H200 at comparable accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.