acceptodds
Under review as a conference paper at ICLR 2027

Unlocking What MLLMs Already See for Hallucination Mitigation

Abstract

Hallucinations remain a critical challenge for Multimodal Large Language Models (MLLMs), which often generate statements inconsistent with the input image. In this work, we uncover a : although an MLLM may hallucinate during generation, it can often recognize these hallucinations when its output is decomposed into atomic claims and each claim is independently verified against the image. Building on this observation, we propose a self-supervised framework called (elf-erification uided n-olicy istillation), which turns this self-verification capability into more reliable generation behavior. SVG-OPD first elicits complementary fact-oriented and probing responses, decomposes them into atomic claims, and self-verifies each claim to identify hallucinated and non-hallucinated content. These verified claims are then provided as privileged information to a frozen teacher, whose token-level distributions supervise student-generated trajectories through on-policy distillation. Unlike existing training methods, SVG-OPD requires neither external knowledge nor a stronger teacher model, instead exploiting the MLLM's own visual verification capability for hallucination-aware post-training. Extensive experiments across generative and discriminative hallucination benchmarks, model families, and scales show that SVG-OPD substantially mitigates hallucinations while broadly preserving general multimodal capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.