acceptodds
Under review as a conference paper at ICLR 2027

Learning to Verify: A Counterfactual-Guided Dual-Loop Agent Framework for Multimodal Hallucination Mitigation

Abstract

Mitigating multimodal hallucination is essential for the reliable utilization of multimodal large language models in image-grounded tasks. Existing approaches perform mitigation by adjusting decoding, strengthening visual information during generation, or using verification to guide response correction. However, verification itself may accept unsupported statements or challenge correct ones, allowing repeated revision to retain hallucinations and remove valid information. In this paper, we propose Learning to Verify (L2V), a dual-loop agent framework for multimodal hallucination mitigation that learns a verifier's criteria through counterfactual analysis of its judgments. The inner loop performs inference by checking a main agent's response against the image and guiding revision until the response is accepted or an uncertainty response is returned. The outer loop operates only during training, using reference responses and counterfactual interventions to identify verification errors whose correction improves the final response. The identified verification errors guide targeted revisions to the verifier's criteria, which are evaluated through held-out replay on examples excluded from criteria refinement and adopted only when they improve response quality over the current criteria. Experiment results and analysis on multimodal hallucination benchmarks show that the proposed L2V outperforms strong baselines and existing approaches, demonstrating its effectiveness in mitigating multimodal hallucination.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.