Who Is the Real Witness? Localizing Attention Heads through Image-Induced Semantic Shifts in LVLMs
Abstract
Large vision-language models (LVLMs) have demonstrated strong performance on multimodal tasks. However, they may overrely on prior knowledge when it conflicts with visual information, producing plausible predictions that are inconsistent with the input image. Previous studies have sought to mitigate this issue by identifying and enhancing attention heads that are sensitive to visual input. We observe that visually sensitive heads do not always capture the key cues needed for the current prediction. In this paper, we propose Witness Identification Through Next-token Evidence of Semantic Shifts (WITNESS), a training-free method. It projects each head's outputs into the vocabulary space and compares the resulting predictive distributions under text-only and image-text conditions over the same candidate token set. The resulting image-induced semantic shifts are quantified using Jensen-Shannon (JS) divergence. WITNESS then selects and enhances heads exhibiting larger shifts and combines this intervention with contrastive decoding. We conduct extensive experiments on the counterfactual datasets WHOOPS-AHA! and SynConFact using four LVLMs. Compared with the base models, WITNESS achieves relative improvements of 6.49% and 16.16% in accuracy averaged across models on the two datasets, respectively. Our code is available at https://anonymous.4open.science/r/WITNESS-57B7.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.