acceptodds
Under review as a conference paper at ICLR 2027

Pairwise-Safe, Jointly-Unsafe: Higher-Order Residual Risk in Parallel Decoding of Diffusion Language Models

Abstract

Diffusion language models (dLLMs) obtain much of their inference advantage by committing multiple masked positions in parallel. Existing fast decoders therefore rely on token confidence or pairwise dependency estimates to decide which positions can be committed together. We identify a blind spot in this design: a token bundle can appear pairwise safe while remaining jointly unsafe. We formalize the portion of joint commitment risk that is unexplained by pairwise dependency, confidence, bundle size, and diffusion stage as Higher-Order Residual Risk (HOR). We first perform a Phase-0 diagnosis on Dream-7B and LLaDA. Pairwise-safe/jointly-unsafe bundles constitute only 13.9–18.2% of observed commit bundles, yet account for 49.2–58.6% of decoding errors, indicating a high-leverage failure mode. We then introduce a lightweight HOR predictor that combines pooled hidden-state representations with pairwise and meta features, and use its prediction as a pre-commit signal to commit, split, or defer candidate bundles. The full predictor reaches 0.896 AUROC and 0.792 R² for joint-risk prediction. Across GSM8K, MATH500, HumanEval+, and MBPP+, HOR-aware decoding improves accuracy/pass@1 over DEMASK by 4.1 points on average while reducing NFE by 8.0% and latency by 10.3%. HOR selection is also complementary to CoCommit-style within-set coordination, yielding further gains. These results suggest that pairwise safety is insufficient for reliable parallel dLLM decoding and that residual set-level risk can be both measurable and actionable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.