Beyond Attention Black Holes: Energy Landscape Guidance for Rectifying Feature Stagnation in VLMs
Abstract
Large Vision-Language Models have achieved remarkable success in multimodal reasoning, yet they remain surprisingly vulnerable to low-level visual distraction. In this work, we identify and formalize a ubiquitous but overlooked phenomenon: Attentional Drift induced by artificial image padding that corresponds to redundant regions in natural scenes. We observe that semantically vacant black pixels and low-entropy background patches often act as attentional black holes, disproportionately absorbing the model's focus and leading to severe reasoning degradation. Our investigation reveals that this drift is triggered by a Dual Trap inherent in the Transformer architecture: 1) a positional bias toward the end of the sequence, and 2) an attention sink induced by feature stagnation, where stable and uninformative tokens act as attention sinks. To mitigate this universal issue, we propose Energy Landscape Guidance (ELG), a training-free, plug-and-play framework that conceptualizes the evolution of hidden states as trajectories within a high-dimensional manifold. ELG introduces an Energy Scoring Module that leverages spectral analysis to capture evolution flux, thereby isolating various forms of pathological tokens without the need for prior masks. This is followed by a Latent Rectification Module that steers hidden states toward high-information semantic basins via gradient-based intervention during inference. Extensive evaluations on VMCBench demonstrate that ELG consistently rehabilitates various VLMs from visual illusions. Notably, Qwen2.5-VL-7B equipped with ELG outperforms 72B-scale competitors and proprietary models like GPT-4o.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.