Confident Thinking: Multi-Layer Entropy-Guided Latent Reasoning for Multimodal Long-Context RAG
Abstract
Multimodal long-context retrieval-augmented generation (MLRAG) is increasingly important for knowledge-intensive question answering, as it enables models to retrieve and reason over extensive textual and visual evidence. However, during answer generation, multimodal large language models remain prone to hallucination when synthesizing retrieved cross-modal evidence, particularly when the evidence is heterogeneous, redundant, or conflicting. Existing methods commonly employ explicit chain-of-thought (CoT) to facilitate evidence reasoning, but its token-by-token generation forces the model to commit to discrete intermediate states, allowing unreliable predictions to propagate through subsequent steps and ultimately yield hallucinated answers. To address this problem, we propose Confident Thinking, a confidence-guided framework that switches from explicit to latent reasoning upon detecting genuine reasoning uncertainty, preserving candidate semantics in continuous space while incorporating retrieved evidence to guide subsequent updates. First, we introduce a multi-layer entropy consensus mechanism that evaluates predictive uncertainty across the last few Transformer layers and triggers latent reasoning only when all these layers exhibit high entropy, reducing unnecessary switches caused by final-layer-specific ambiguity. Second, we develop an evidence-calibrated soft-thinking mechanism that dynamically selects question- and state-relevant multimodal evidence and injects it into soft tokens through uncertainty-aware gating, constraining recursive latent updates to mitigate reasoning drift and error accumulation. Extensive experiments on multiple multimodal long-context question-answering benchmarks demonstrate that Confident Thinking consistently improves answer accuracy and overall performance across different settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.