Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering
Abstract
Large vision–language models (LVLMs) have achieved strong performance in visual reasoning and understanding, but can still produce hallucinations, generating content that is not adequately supported by the visual input. Many existing methods focus on the language side by intervening in model outputs or decoding processes, while the role of visual representations remains less explored. We therefore investigate hallucinations in the model's internal visual representation space. We first use sparse autoencoders (SAEs) to transform dense and entangled visual representations into a sparse, interpretable feature space. We identify two types of SAE features: always-on neurons that recur across images and image-specific neurons that respond to the content of the current image. Intervention experiments reveal that hallucinations are associated with the spurious activation of neurons unrelated to the input image. Further experiments show that enhancing image-relevant neurons while suppressing irrelevant ones can mitigate hallucinations. Based on these findings, we propose Contrastive Neuron Steering (CNS), which uses contrastive feature steering to correct spurious visual activations. CNS further applies Always-on Neuron Compensation (ANC) to offset changes in activation mass and maintain stable token-level representations. Experiments on hallucination and general-purpose benchmarks show that CNS reduces hallucinations and improves response faithfulness to the input image without compromising general multimodal performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.