Unlocking What MLLMs Already See for Hallucination Mitigation
Abstract
Hallucinations remain a critical challenge for Multimodal Large Language Models (MLLMs), which often generate statements inconsistent with the input image. In this work, we uncover a : although an MLLM may hallucinate during generation, it can often recognize these hallucinations when its output is decomposed into atomic claims and each claim is independently verified against the image. Building on this observation, we propose a self-supervised framework called (elf-erification uided n-olicy istillation), which turns this self-verification capability into more reliable generation behavior. SVG-OPD first elicits complementary fact-oriented and probing responses, decomposes them into atomic claims, and self-verifies each claim to identify hallucinated and non-hallucinated content. These verified claims are then provided as privileged information to a frozen teacher, whose token-level distributions supervise student-generated trajectories through on-policy distillation. Unlike existing training methods, SVG-OPD requires neither external knowledge nor a stronger teacher model, instead exploiting the MLLM's own visual verification capability for hallucination-aware post-training. Extensive experiments across generative and discriminative hallucination benchmarks, model families, and scales show that SVG-OPD substantially mitigates hallucinations while broadly preserving general multimodal capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.