TEMPO: Temporal and Positional Precision for Diffusion Language Model Agents
Abstract
Diffusion large language model (dLLM) agents recompute a fixed-length generation region at each denoising step. Reserved positions beyond the emitted call account for most of this computation. Removing them reduces accuracy because they provide context for uncommitted predictions, requiring a reduction in computation per position. Quantisation reduces this cost, but existing quantisers use a fixed configuration across steps and positions, irrespective of the decoding state defined by committed tokens and current predictions. We identify three phenomena that enable precision to adapt to this state. Filler-first ordering places end-of-text and padding commitments before content commitments, allowing early steps to run at low precision. Margin persistence correlates logit gaps across steps, indicating the risk of quantisation-induced prediction changes before each pass. Slack tolerance permits low-precision updates of the reserved region with limited changes in accuracy. These phenomena motivate TEMPO (Temporal and Positional Orchestration of Precision), a training-free hierarchical precision schedule. A step-wise gate selects low-precision steps using cached gaps. On other steps, position-wise routing retains full precision only for the prompt and estimated content region. Across four agent benchmarks and three backbones, TEMPO achieves 84% mean call fidelity, compared with 76% for full quantisation and 66% for Learn2PD. Adding TEMPO to parallel decoding with boundary fill yields a further to speedup on Spider and API-Bank. The full composition achieves up to at 128 denoising steps. Our code is openly accessible to support future research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.