acceptodds
Under review as a conference paper at ICLR 2027

Distilling Language Knowledge Across Generative Processes

Abstract

Diffusion language models enable parallel text generation. Prior work has shown that distillation from pretrained autoregressive (AR) models offers a promising route to acquiring strong language capabilities. However, the causal-prefix supervision used in prior work does not fully account for the mismatch between the two generation processes: an AR teacher predicts from a left prefix, while a diffusion student reconstructs missing text using observed context on both sides. To address this mismatch, we introduce CrossDLM, a framework for distilling language knowledge across these generative processes. CrossDLM uses the teacher's sequence-level scores to evaluate candidate tokens in the context of subsequent text, providing supervision for bidirectional reconstruction. This supervision incorporates visible evidence while accounting for uncertainty in unresolved content. Shared teacher computation across token alternatives and aggregation over a small number of student-generated completions make this supervision practical. With on-policy self-distillation, CrossDLM converts pretrained AR checkpoints into DLMs using approximately 0.10B training tokens per model. It outperforms causal-prefix distillation across a range of benchmarks and model families.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.