Distilling Language Knowledge Across Generative Processes
Abstract
Diffusion language models enable parallel text generation. Prior work has shown that distillation from pretrained autoregressive (AR) models offers a promising route to acquiring strong language capabilities. However, the causal-prefix supervision used in prior work does not fully account for the mismatch between the two generation processes: an AR teacher predicts from a left prefix, while a diffusion student reconstructs missing text using observed context on both sides. To address this mismatch, we introduce CrossDLM, a framework for distilling language knowledge across these generative processes. CrossDLM uses the teacher's sequence-level scores to evaluate candidate tokens in the context of subsequent text, providing supervision for bidirectional reconstruction. This supervision incorporates visible evidence while accounting for uncertainty in unresolved content. Shared teacher computation across token alternatives and aggregation over a small number of student-generated completions make this supervision practical. With on-policy self-distillation, CrossDLM converts pretrained AR checkpoints into DLMs using approximately 0.10B training tokens per model. It outperforms causal-prefix distillation across a range of benchmarks and model families.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.