HalluMirror: Discovering and Mitigating Multimodal Hallucinations via Co-Evolutionary Self-Play
Abstract
Mitigating hallucinations in Multimodal Large Language Models (MLLMs) increasingly relies on constructed training supervision. However, existing methods typically optimize against predefined hallucination patterns, curated preferences, or model-generated samples whose challenges are not explicitly adapted to the model's evolving failure modes. We identify this mismatch as a key limitation: as the model improves, effective supervision should continually shift toward its current unresolved hallucination weaknesses. To this end, we propose HalluMirror, a co-evolutionary adversarial self-play framework that continually transforms model failures into new training challenges. HalluMirror couples an Inducer, optimized to generate fine-grained existence, attribute, and relation hallucinations that expose the current Challenger's weaknesses, with a Challenger that learns to detect and repair these evolving hard cases. Through alternating optimization, weakness discovery, adversarial exposure, and targeted repair form a closed self-improvement loop. To reliably scale this process, we further introduce a training-free grounding-based hallucination verifier that filters self-play-generated hallucinations without hallucination-specific human annotations. Extensive experiments across multiple MLLMs and hallucination benchmarks show that HalluMirror consistently improves robustness against diverse hallucination types. On GLM-4.1V-9B, HalluMirror improves performance across all eight hallucination benchmarks, with co-evolution identifying and addressing the model’s vulnerability to relation hallucinations and yielding a 14.5% relative gain on FINER’s multi-relation subset. All data and code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.