acceptodds
Under review as a conference paper at ICLR 2027

Meta-CogGran: Reasoning as Online Latent Cognitive State Evolution for MLLMs

Abstract

Recent progress in multimodal large language models (MLLMs) has largely depended on CLIP-based visual encoders. While these models excel at capturing high-level semantic alignment, they often struggle to preserve fine-grained visual evidence. By comparison, the recent DINOv3 has exhibited emergent capabilities in fine-grained pixel-level perception, but still lacks the higher-level semantic abstraction required for alignment with LLMs and complex multimodal reasoning. These complementary limitations reveal a fundamental bottleneck: existing MLLMs lack a mechanism to dynamically reconcile fine-grained perception with high-level semantics during reasoning. We argue that the key challenge lies not merely in obtaining richer visual representations, but in enabling perceptual evidence and semantic knowledge to continuously interact, update, and converge during reasoning. This raises a central question: How can fine-grained perception and high-level semantics be dynamically integrated into evolving cognitive states for multimodal reasoning? To address this challenge, we propose **Meta-CogGran**, which formulates multimodal reasoning as online latent cognitive state evolution. It iteratively updates structured cognitive states through interactions with visual evidence and semantic memory, while a semantic equilibrium mechanism terminates reasoning once the state becomes stable and self-consistent. Extensive experiments demonstrate approximately **45% ↑** improvement in accuracy and **30% ↓** reduction in hallucination. Code is available at [Meta-CogGran](https://anonymous.4open.science/r/Meta-CogGran-5331/).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.