acceptodds
Under review as a conference paper at ICLR 2027

Temporal-Aware Trajectory Self-Distillation for Fast and Accurate Diffusion LLM

Abstract

Trajectory distillation is a promising training-based approach for accelerating diffusion large language models (dLLMs) by adapting them to generate more tokens per forward pass. However, existing methods construct trajectories from the base model alone and supervise masked tokens without distinguishing their decoding horizons, which limits the knowledge transferred through distillation and yields a suboptimal accuracy–parallelism trade-off. Enriching the knowledge carried by these trajectories and tailoring supervision to temporal decoding dynamics remain open challenges. Inspired by learning with privileged information, and motivated by our observation that the teacher–student distributional gap varies systematically with a token's time to reveal, we propose TAD, a Temporal-Aware trajectory self-Distillation framework. TAD conditions a frozen teacher on target responses available only during training to construct informative decoding trajectories. It then partitions masked tokens into near and distant subsets according to their temporal distance to decoding, applying hard cross-entropy supervision to near tokens for confident parallel decoding and soft distribution matching to distant tokens to avoid premature decisions while preserving future-planning knowledge. Across mathematical reasoning and code generation benchmarks, TAD improves the average accuracy of LLaDA-8B-Instruct from 46.1% to 51.5% and Dream from 58.9% to 61.2%. Its speed-oriented configuration further achieves the highest average Accuracy Under Parallelism scores of 256.6 and 193.9 on the two backbones respectively, which validate that TAD establishes a superior accuracy-parallelism frontier, simultaneously enhancing task accuracy and decoding throughput. Our code is available at https://anonymous.4open.science/r/TAD-B9D0/https://anonymous.4open.science/r/TAD-B9D0/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.