Multi-Hop Attention Graph Sampling for Masked Diffusion LLMs
Abstract
Masked Diffusion LLMs can accelerate language generation through parallel decoding, but unmasking incompatible token combinations can degrade quality. Recent attention-based dependency-aware samplers construct a graph where the masked positions are nodes and the edges correspond to direct attention map weights. Though achieving impressive acceleration, these consider only direct or single-hop relations in the graph, disregarding multi-hop connectivity. We introduce GRID - *Graph Resistance Independent Decoding* - a training-free sampler that replaces the direct-edge coupling score with *effective resistance*, a parameter-free connectivity measure that aggregates both direct and indirect routes between masked positions. To analyze how samplers perform on complex multi-hop graph structures, we introduce *Dep-Bench*, a controlled benchmark with five datasets spanning four graph structures and known dependencies at varying depths. On Dep-Bench, GRID's advantage over direct-edge methods grows consistently with depth. On ParallelBench, a public benchmark focused on parallel decoding, GRID achieves up to decoding-step speedup over the fastest competing sampler, for a given quality threshold of % of the category-level top-1 score. Finally, across 16 leading model and LLM reasoning benchmark combinations, GRID achieves speedup over top-1 decoding and % improvement over the fastest competing sampler, at comparable accuracy. The combined results validate that multi-hop connectivity is a valuable and underutilized signal for dependency-aware parallel decoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.