GROUNDING BEFORE SPEAKING: CROSS-MODAL REFLECTIVE ROUTING FOR HALLUCINATION-RESILIENT MULTIMODAL LARGE LANGUAGE MODELS
Abstract
Multimodal large language models (MLLMs) have rapidly improved on visual question answering, multimodal dialogue, document understanding, and scien tific reasoning, yet they remain vulnerable to hallucination: fluent outputs that are weakly supported by visual evidence. Existing mitigation methods often rely on retrieval, expensive preference optimization, or post-hoc checking, leav ing the autoregressive decoding process itself insufficiently grounded. We pro pose Cross-Modal Reflective Routing (CMRR), a lightweight framework that estimates token-level grounding uncertainty, dynamically routes computation be tween visual and textual streams, and verifies candidate tokens before commit ment. CMRR introduces a reflective uncertainty estimator, a differentiable cross modal router, and a semantic verifier trained with contrastive grounding super vision. The resulting model suppresses visually unsupported tokens while pre serving linguistic fluency. We present a complete training objective, decoding algorithm, and reproducibility package with implementation, visualizations, and evaluation scripts. In pilot benchmark tables prepared for reproducible experi mentation, CMRR-style routing improves hallucination-sensitive metrics across MMHalBench, POPE, HallusionBench, ScienceQA, and MMBench while adding modest decoding overhead. The paper highlights a practical direction for control lable, uncertainty-aware MLLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.