acceptodds
Under review as a conference paper at ICLR 2027

Beyond Global Thresholds: Token-Level Calibrated Adaptive Parallel Decoding for Diffusion Language Models

Abstract

Diffusion language models (DLMs) are a promising alternative to autoregressive models because they can generate multiple tokens in parallel. However, they are still slow in practice since they need many denoising steps. Recent methods speed up inference by using a fixed confidence threshold to decide which tokens can be decoded in parallel. These methods assume that confidence has the same meaning for all tokens, which often does not hold. In this work, we show that different tokens have different reliability even under the same confidence level, which reveals a systematic token-level miscalibration problem in confidence-aware decoding. We also find that confidence is related to final correctness, but the relationship varies across tokens. Based on these observations, we propose Token-Level Calibrated Adaptive Parallel Decoding (TC-APD), which estimates a separate threshold for each token from offline calibration data. During inference, we use these thresholds to guide adaptive parallel decoding without adding extra model components. Experiments on mathematical reasoning and code generation tasks show that our method improves the speed-quality trade-off over fixed-threshold baselines, while keeping generation quality. Overall, we demonstrate the effectiveness and potential of calibrating confidence thresholds at the token level for accelerating diffusion language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.