DM-OPD: Aligning Diffusion Language Models with Autoregressive Teachers
Abstract
Converting strong autoregressive (AR) models into diffusion language models promises parallel generation while retaining AR capabilities. Existing conversion methods fit intermediate predictions, but improving this fit need not improve the distribution of generated text. We analytically demonstrate this mismatch even on the student's own generation trajectories. We introduce Distribution Matching OPD (DM-OPD), which defines distillation on the text distribution induced by the student's actual decoder and derives an unbiased on-policy gradient estimator for each completed-block KL at student-generated prefixes. We make this objective trainable by estimating student text likelihoods, decomposing completed-text feedback into dependency-based action credit, and defining trainable probabilities for refinement decisions. Using Qwen3 initialization and teacher checkpoints with OPDLM's prompt corpus, we distill without answer-correctness rewards. Our 1.7B student outperforms the diffusion students SDAR and OPDLM on all eight benchmarks at 4K and 8K response limits, as does our 4B student except on HumanEval at 4K. With the same decoder and the guard disabled, our objective improves monitoring accuracy by 5.4 points over OPDLM's objective, whose loss keeps falling. Students trained for fewer refinement iterations outperform released diffusion models decoded with the same iteration budget on most benchmarks. These results establish completed-text distribution matching as an effective foundation for AR-to-diffusion distillation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.