acceptodds
Under review as a conference paper at ICLR 2027

HalluGuard: Decoupled Automatic Hallucination Detection and Mitigation in MLLMs

Abstract

Multimodal Large Language Models (MLLMs) often exhibit severe hallucinations, where generated text conflicts with visual inputs. Existing evaluation methods heavily rely on costly human annotators or black-box APIs (e.g., GPT-4o). However, these end-to-end evaluators inherently suffer from multimodal entanglement, making them prone to hallucinations themselves and unreliable as judges. To address this, we introduce HalluGuard, a decoupled and interpretable framework for hallucination evaluation and mitigation. Instead of relying on black-box MLLMs, our AutoEval module decomposes the complex hallucination detection process into atomic visual perception tasks (e.g., grounding and segmentation). This white-box paradigm provides accurate, fine-grained evaluation for attributes prone to hallucination, such as object quantity, color, and relative position. Based on this deterministic feedback, we further propose Interpretable Hallucination Mitigation (IHM), an inference-time intervention mechanism that effectively corrects hallucinations without requiring any parameter updates or preference data. Furthermore, leveraging our fine-grained metric, we reveal a critical “Goodhart's Law” phenomenon in current MLLM alignment: models often hack hallucination metrics by generating shorter, evasive responses rather than genuinely correcting visual-textual misalignments. Extensive experiments demonstrate that HalluGuard achieves high human agreement in evaluation and effectively mitigates hallucinations, offering a robust alternative to current end-to-end paradigms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.