VIVA: Visual Information Dispersion via Attention Guidance for Parallel Decoding in Diffusion MLLMs
Abstract
Diffusion multimodal large language models (dMLLMs) enable parallel decoding by predicting multiple masked tokens using bidirectional context, offering a promising alternative to autoregressive MLLMs. However, existing token selection strategies largely determine which tokens to decode based on prediction confidence at each step, without considering how visual conditioning influences the overall decoding trajectory. In this work, we analyze dMLLM decoding trajectories and identify two key issues: visually relevant tokens tend to be decoded in later stages, limiting their contribution as context for earlier predictions, and tokens describing visual content often influence each other's predictions, making their simultaneous decoding prone to prediction errors. Based on these observations, we propose **V**isual **I**nformation Dispersion **v**ia **A**ttention Guidance (VIVA), an inference-time decoding method that disperses visually informative tokens both across decoding steps and within each step. VIVA uses token-to-image attention to distribute visually relevant tokens across the decoding trajectory and limits their simultaneous selection within each step, while preserving the ability to decode tokens in parallel. Across six benchmarks and three open-source dMLLMs, VIVA demonstrates strong performance against existing decoding baselines, particularly when multiple tokens are decoded in parallel. Specifically, on LaViDa-Reason, VIVA improves CIDEr from 67.5 to 86.0 on COCO Captioning and MMBench accuracy from 43.0 to 51.4 over confidence-based decoding when decoding eight tokens per step. Further analyses and qualitative examples confirm that VIVA effectively mitigates both identified decoding issues.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.