Fine-Tuning Discrete Diffusion LLMs by Distilling from Autoregressive Teachers
Abstract
Discrete diffusion language models (dLLMs) have recently emerged as an alternative paradigm for text generation, with the potential to achieve more efficient inference than autoregressive (AR) models. However, AR models still continue to achieve stronger generation quality than dLLMs. The AR models provide strong token-level predictive distributions and benefit from more established training methods. We explore how to leverage AR models during training while preserving the parallel generation capability of dLLMs at inference time. However, directly distilling AR knowledge into dLLMs is challenging because AR models and dLLMs condition on different contexts, leading to mismatched predictive distributions. To address this challenge, we propose Weight-annEaled Autoregressive distillatioN for dLLMs (WEAN), the first AR-to-dLLM distillation method designed for downstream fine-tuning. WEAN controls where and when AR supervision is applied: it selectively transfers AR knowledge to reliable token positions and dynamically adjusts the distillation strength according to denoising steps and training stages. Specifically, WEAN focuses distillation on structural tokens, where AR and dLLM predictions are more compatible, and gradually reduces AR guidance as the dLLM adapts to the target distribution. Experimental results show that WEAN consistently improves both generation quality and downstream utility compared with the strongest baselines by 17.0% and 2.03% on studied text datasets. WEAN generalizes to code and mathematics domains, showing its broad applicability. We release the replication package on the anonymous link.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.