ACE: Attention Coalition Envelopes for Parallel Decoding in Diffusion Language Models
Abstract
Masked diffusion large language models (dLLMs) enable parallel generation by predicting multiple unresolved positions in a single denoiser pass, but individually confident predictions need not be jointly compatible. We study this tension through a cooperative-game view of parallel commitment. Under a coherent reference distribution, the KL risk of a parallel commit decomposes into individual prediction errors and a supermodular dependence cost. This structure leads to an ordered modular envelope that assigns each candidate a fixed charge before selection, reducing coalition formation to sorting and a cumulative sum. We instantiate this principle as the Attention Coalition Envelope (ACE), which uses native attention to form compatibility-aware token coalitions without training or additional denoiser evaluations; its formal guarantee applies to the attention-induced surrogate game. Across six benchmarks on LLaDA and Dream, ACE reaches the empirical quality–efficiency Pareto frontier in all 12 benchmark–backbone settings, with 30/36 evaluated operating points being Pareto-optimal. It also attains the highest measured accuracy in 10/12 settings, including 5 statistically supported leads. On two matched LLaDA reasoning controls, ACE reduces denoiser evaluations by 32–36% at comparable accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.