acceptodds
Under review as a conference paper at ICLR 2027

CogRef: Cognitive Reflection for Hallucination Mitigation in Multimodal Large Language Models

Abstract

Hallucination remains a major challenge for multimodal large language models (MLLMs). Existing mitigation methods often modify the generation process of MLLMs or rely on external verification tools, potentially constraining the original capabilities of the base MLLM. Motivated by human cognitive learning, where past experiences are reflected upon and reused to guide subsequent behavior, we propose CogRef, a cognitive reflection framework that decouples multimodal generation from hallucination supervision. A frozen MLLM serves as the cognitive core, while a learned supervisory circuit evaluates its response and provides structured feedback for final generation. To train the circuit, we transform past response cases into reusable multimodal experience, and internalize it through Concordance-Calibrated On-Policy Self-Distillation (CC-OPSD). CC-OPSD coordinates experience- and evidence-privileged teachers, enabling the supervisory circuit to internalize both response-handling experience and fine-grained visual grounding capability. At inference, CogRef requires neither privileged information nor external verification tools. Experiments across diverse MLLM families and multimodal benchmarks demonstrate that CogRef consistently mitigates hallucinations while preserving general multimodal capability, with the largest improvement reaching points (MHaluBench: ). Moreover, the supervisory circuit generalizes across different MLLMs without model-specific retraining, while selectively correcting hallucinated content and preserving visually grounded responses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.