acceptodds
Under review as a conference paper at ICLR 2027

Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering

Abstract

Large vision–language models (LVLMs) have achieved strong performance in visual reasoning and understanding, but can still produce hallucinations, generating content that is not adequately supported by the visual input. Many existing methods focus on the language side by intervening in model outputs or decoding processes, while the role of visual representations remains less explored. We therefore investigate hallucinations in the model's internal visual representation space. We first use sparse autoencoders (SAEs) to transform dense and entangled visual representations into a sparse, interpretable feature space. We identify two types of SAE features: always-on neurons that recur across images and image-specific neurons that respond to the content of the current image. Intervention experiments reveal that hallucinations are associated with the spurious activation of neurons unrelated to the input image. Further experiments show that enhancing image-relevant neurons while suppressing irrelevant ones can mitigate hallucinations. Based on these findings, we propose Contrastive Neuron Steering (CNS), which uses contrastive feature steering to correct spurious visual activations. CNS further applies Always-on Neuron Compensation (ANC) to offset changes in activation mass and maintain stable token-level representations. Experiments on hallucination and general-purpose benchmarks show that CNS reduces hallucinations and improves response faithfulness to the input image without compromising general multimodal performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.