acceptodds
Under review as a conference paper at ICLR 2027

Attend to Vision When It Matters: Visual Dependency-Aware Attention Rebalancing for Large Vision-Language Models

Abstract

Large Vision-Language Models (LVLMs) commonly suffer from visual hallucinations, generating content that deviates from visual evidence. This problem stems from the dominance of linguistic priors in autoregressive generation and the insufficient visual grounding. Existing inference-time attention interventions mitigate hallucinations by strengthening visual attention, but typically rely on native attention patterns and adopt static or uniform enhancement strategies, failing to distinguish informative visual tokens, identify layers that benefit most from intervention, or adapt to the dynamic demand for visual evidence during generation. Such indiscriminate intervention may alleviate hallucinations at the cost of language generation quality. To address this, we propose ViDRA, a fine-grained attention rebalancing framework. During prefill, ViDRA estimates visual token importance by jointly considering representation stability and instruction relevance, while deriving layer-wise intervention strength from visual activity across Transformer layers. During decoding, a lightweight Visual Dependency Gate predicts the visual dependency of the upcoming token, enabling visual attention to be enhanced only when needed. Extensive experiments on visual hallucination benchmarks and general vision-language tasks show that ViDRA effectively reduces visual hallucinations while preserving language generation quality and inference efficiency, achieving state-of-the-art performance across multiple LVLM backbones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.