acceptodds
Under review as a conference paper at ICLR 2027

Learning What to Align: Automatic Target Discovery for Hallucination Mitigation in MLLMs

Abstract

Multimodal large language models often suffer from visual hallucination, where generated responses are inconsistent with the input image. Existing preference-based alignment methods mainly rely on sequence-level supervision, making it difficult to localize visual evidence and response tokens responsible for hallucinated predictions. Recent target-level approaches improve alignment precision but usually depend on manually annotated regions or token spans. In this paper, we propose Learning What to Align with Target Discovery (LWA-TD), which leverages minimal visual contrast to identify visual evidence and response tokens responsible for grounded semantic differences. Given counterfactual image-response pairs, LWA-TD identifies question-relevant visual changes and estimates token sensitivity to counterfactual visual changes. The discovered targets are further incorporated into preference optimization and grounding regularization to improve visual-text consistency. Experiments on multiple hallucination benchmarks demonstrate that LWA-TD effectively reduces hallucination across different model scales and backbones, while maintaining strong general multimodal reasoning

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.