RAPD: Repulsion-Attraction Energy for Parallel Decoding of Diffusion Large Language Models
Abstract
Parallel decoding accelerates masked diffusion large language models by committing multiple positions per forward pass, but selecting a set of mutually compatible positions remains challenging. Existing methods exploit inter-position cues such as locality and dependency, but most of them model these interactions in only one direction, either promoting or suppressing joint commitment. We propose repulsion-attraction parallel decoding (RAPD), a training-free strategy that combines individual commitment evidence and signed pairwise interactions in a single energy function for commit-set selection. Nearby positions attract each other and promote joint commitment, whereas positions that compete for the same candidate tokens repel each other and suppress conflicting commitment. RAPD iteratively refines commitment probabilities using variational updates, without additional model forward passes. On LLaDA-8B-Instruct, RAPD uses a single configuration across GSM8K, MATH, HumanEval, and MBPP, achieving the fewest forward passes among the compared methods while maintaining competitive accuracy. On HumanEval, RAPD matches the pass@1 of Fast-dLLM with 27.8% fewer forward passes and achieves 4.73x the throughput of sequential decoding. On the same benchmark, threshold sweeps further show that RAPD matches or exceeds the accuracy of Fast-dLLM and LocalLeap with fewer forward passes across all tested baseline thresholds. Our analysis shows that attraction reduces decoding steps, while repulsion mitigates the accuracy loss from attraction-only decoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.