TreePC: Posterior-Consistent and Dependency Aware Decoding for Few-Step Diffusion Language Models
Abstract
Diffusion language models (dLLMs) enable parallel generation by unmasking multiple tokens at each denoising step. Existing dLLM efficient decoding methods usually achieve speedups by substantially reducing decoding steps. However, aggressive step reduction forces dLLM to commit far more larger token sets at each step, thereby degrading generation quality. We identify two primary causes of quality degradation in few-step dLLM efficient decoding: (1) ***marginal drift***, where a few-step student’s per-position posteriors deviate from those of a fine-step teacher; and (2) ***dependency error***, where independent sampling ignores compatibility among tokens committed within the same step. To fill the identified research gap, we introduce **TreePC**, a posterior-consistent and dependency-aware decoding framework that first aligns student marginals through Posterior Consistency LoRA (PC-LoRA), then predicts pairwise dependencies from teacher counterfactuals, constructs a counterfactual-dependency maximum spanning tree, and performs parent-conditioned corrections along the tree. A learnable global scale jointly controls the correction budget with edge-specific gates, allowing TreePC to exploit informative dependencies while preserving calibrated posteriors when conditional evidence is weak. During inference, TreePC invokes dLLM backbone model exactly once per diffusion step; its dependency estimator, tree construction, and correction module are lightweight and leave backbone neural function evaluations unchanged. Experimental results on GSM8K, MATH-500, HumanEval, and MBPP demonstrate that TreePC improves task accuracy and reduces marginal and conditional KL divergence while bringing very modest decoding latency. These results establish structured posterior correction as an effective quality-efficiency trade-off for few-step dLLM generation. Code is available at: [TreePC](https://anonymous.4open.science/r/TreePC_efficient_dLLM_decoding-C22F/).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.