acceptodds
Under review as a conference paper at ICLR 2027

Reducing Hallucinations in Multimodal LLMs with Modality-Aware Sparse Autoencoders

Abstract

Multimodal large language models (MLLMs) often hallucinate visual content absent from the input image. Preference optimization has become a common post-training approach for mitigating hallucinations, but existing preference signals often rely on LLM-as-Judge supervision that can be sensitive to linguistic style rather than visual grounding. We thus introduce Modality-Aware Sparse Autoencoders (MA-SAE), an SAE architecture that separates vision-specific, text-specific, and cross-modal features to provide a visual grounding signal for preference labeling. By isolating modality-specific patterns, such as hedging and stylistic scaffolding, MA-SAE learns interpretable cross-modal features from which we derive a grounding score measuring how well a response is supported by the image. Across three MLLM backbones, MA-SAE learns more coherent and better aligned cross-modal features than existing SAE baselines, achieving the best performance across five feature quality metrics. Using its grounding score for preference labeling, MA-SAE substantially reduces hallucinations under Direct Preference Optimization (DPO), consistently outperforming LLM-as-Judge and other preference-labeling methods across three benchmarks. The grounding signal further transfers to mDPO and on-policy preference tuning, where MA-SAE achieves strong overall performance. These results show that interpretable cross-modal features can provide a direct signal of visual grounding, offering an alternative to LLM-as-Judge supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.