CARE-Flow: Mitigating Inter-Image Hallucination in Multi-Image MLLMs
Abstract
Hallucination remains a major reliability challenge for multimodal large language models (MLLMs), and becomes more complex when reasoning across multiple images. Even when visual content remains grounded, models may select irrelevant evidence, omit required evidence, or incorrectly combine supported evidence. We call such failures in using grounded evidence across images inter-image hallucinations. Standard multimodal inference entangles determining what evidence a question requires, acquiring that evidence from the images, and forming the response, making it difficult to establish which visual information actually drives the inference. Because observed association cannot establish such dependence, the use of evidence must be tested through intervention. We propose CARE-Flow, a causal framework that decouples evidence specification and acquisition from response formation through an explicit evidence mediator. This mediator makes image-specific contributions independently testable before response generation and constrains downstream inference to the established evidence, yielding a constructed conditional front-door structure for tracing visual influence. Existing multi-image benchmarks mainly score final answers and provide limited diagnosis of how evidence use fails. We therefore introduce CARE-Bench, a 12,000-question benchmark that diagnoses failures in evidence formation and combination across multiple visual domains. Across three MLLMs and seven multi-image benchmarks, CARE-Flow consistently improves accuracy, outperforming the strongest baselines by 4.28–7.53 percentage points on CARE-Bench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.