acceptodds
Under review as a conference paper at ICLR 2027

Different Factorizations, Same Path: On-Policy Distillation from Autoregressive to Diffusion Language Models

Abstract

Diffusion language models (DLMs) have drawn increasing attention for their parallel decoding capability, but still lag behind strong autoregressive language models (ARLMs) in generation quality. This motivates the use of on-policy distillation to transfer the stronger generation capability of ARLMs to DLMs. Yet, such cross-architecture distillation is challenging because ARLMs and DLMs employ fundamentally different conditioning structures for token prediction. To bridge this mismatch, we augment the ARLM response distribution with a prescribed forward masking process to construct a teacher distribution over the same path space as the DLM decoding process. We then introduce Path KL to align the student and teacher path distributions, and prove that it upper-bounds the reverse KL between their final response distributions, enabling teacher evaluation of complete-response to guide the student denoising trajectories. Experiments demonstrate consistent improvement across model scales with average benchmark gains of up to 3.15 points, which become more pronounced under highly parallel decoding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.