ARA: Asynchronous Restoration Attribution for Attention Heads in Diffusion Language Models
Abstract
Diffusion large language models (DLLMs) reconstruct masked tokens through successive denoising stages, yet the importance of their attention heads throughout this process remains unclear. We study how head importance changes with the available context, using masking rate as a controlled proxy for denoising stage. Head rankings retain shared structure but generally become less similar as masking rates move farther apart, revealing structured variation that a single global ranking can obscure. To quantify head contributions under these different reconstruction conditions, we introduce Asynchronous Restoration Attribution (ARA). The method averages contributions over paths that gradually restore suppressed head outputs according to independently sampled schedules. Unlike synchronized restoration, it evaluates each head while others are restored to different degrees. This allows it to capture interaction-dependent contributions that a synchronized path can miss. Experiments on two DLLM families across eight tasks support the effectiveness of ARA in identifying functionally important heads through dynamic pruning and guiding adaptive sparse attention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.