acceptodds
Under review as a conference paper at ICLR 2027

MARVEL: Multi-Agent Reinforced Visual Examination for Hallucination

Abstract

Hallucination remains a critical challenge for vision-language models (VLMs), especially in real-world scenarios where models generate image-inconsistent content. Existing mitigation methods typically rely on human annotations or external supervision, limiting their scalability to large-scale unlabeled image corpora. We propose MARVEL, a multi-agent framework that mitigates hallucination in VLMs via reinforced visual examination, without any ground-truth labels. MARVEL casts factual verification as a multi-agent examination process: model responses are decomposed into atomic claims, reformulated as visual questions, and examined against the image to produce consistency signals that serve as reinforcement learning rewards. To improve reward reliability, we further introduce a Visual Claim Regulation (VCR) mechanism that encourages informative, image-grounded descriptions while suppressing trivial or redundant claims. We evaluate MARVEL on 12 benchmarks spanning hallucination evaluation (+3.30), real-world VQA (+1.74), and general multimodal understanding benchmarks (+4.07), where it consistently outperforms existing methods without human-annotated factual supervision. Moreover, MARVEL generalizes effectively across model scales, demonstrating its potential as a scalable hallucination mitigation solution on large-scale real-world unlabeled image corpora.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.