Regulating Dynamic Attention Allocation in Diffusion Multimodal Large Language Models
Abstract
Diffusion multimodal large language models (dMLLMs) generate text by iteratively denoising a masked response, and at every step, image and output tokens compete for the same softmax attention budget. We find that a small set of output tokens, which we term output sinks, absorbs a disproportionate share of this budget. These sinks shift across positions between steps and claim a growing share of attention as denoising proceeds, at the expense of the image. Such diversion limits the visual evidence available for generation and can induce hallucination. Most sinks are low-semantic tokens (masks, punctuation, and function words), but some are content words. Motivated by these findings, we propose Elbow Proportional Suppression (EPS), a training-free attention intervention. For each head at each step, EPS detects sinks with a parameter-free elbow criterion and attenuates only the low-semantic ones relative to an adaptively selected anchor token via an additive logit shift. We prove that EPS redistributes the suppressed attention proportionally to all remaining keys; empirically, most of it is reallocated to the image. Controlled interventions show that both visual reallocation and the choice of suppressed tokens reduce hallucination. Across three dMLLMs, EPS reduces object hallucination (by up to 8.0 points in CHAIR) and improves description quality and multimodal reasoning, achieving the best results on 24 of 26 metrics against five training-free baselines. Code and analysis tools will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.