acceptodds
Under review as a conference paper at ICLR 2027

SensiEdit: Sensitivity-Guided Spatial Attention Editing for Hallucination Mitigation in LVLMs

Abstract

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal reasoning and complex scene understanding. However, they still suffer from hallucination, generating descriptions that contradict visual facts. Attention heads differ in their sensitivity to image content, motivating selective modulation of visual attention. To exploit this variation, we propose SensiEdit, a lightweight framework for sensitivity-guided spatial attention editing. Specifically, we introduce Attention-based Visual Sensitivity (AVS) to characterize each head’s response to changes in visual input and train a lightweight predictor to estimate this signal. The predicted sensitivity guides stronger visual attention modulation for less sensitive heads, while learnable spatial biases determine the editing strength at different visual positions. The LVLM backbone remains frozen throughout training, and the spatial biases are learned from only 3K images. Experiments on LLaVA-1.5-7B and Qwen2.5-VL-7B, together with supplementary results on three additional backbones, demonstrate the effectiveness of SensiEdit in mitigating object hallucination. On LLaVA-1.5-7B, SensiEdit achieves relative reductions of 56% and 68% in CHAIR_S and CHAIR_I, respectively, compared with the unedited baseline, while maintaining performance on the evaluated general-capability benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.