Attend to Vision When It Matters: Visual Dependency-Aware Attention Rebalancing for Large Vision-Language Models
Abstract
Large Vision-Language Models (LVLMs) commonly suffer from visual hallucinations, generating content that deviates from visual evidence. This problem stems from the dominance of linguistic priors in autoregressive generation and the insufficient visual grounding. Existing inference-time attention interventions mitigate hallucinations by strengthening visual attention, but typically rely on native attention patterns and adopt static or uniform enhancement strategies, failing to distinguish informative visual tokens, identify layers that benefit most from intervention, or adapt to the dynamic demand for visual evidence during generation. Such indiscriminate intervention may alleviate hallucinations at the cost of language generation quality. To address this, we propose ViDRA, a fine-grained attention rebalancing framework. During prefill, ViDRA estimates visual token importance by jointly considering representation stability and instruction relevance, while deriving layer-wise intervention strength from visual activity across Transformer layers. During decoding, a lightweight Visual Dependency Gate predicts the visual dependency of the upcoming token, enabling visual attention to be enhanced only when needed. Extensive experiments on visual hallucination benchmarks and general vision-language tasks show that ViDRA effectively reduces visual hallucinations while preserving language generation quality and inference efficiency, achieving state-of-the-art performance across multiple LVLM backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.