Thought–State Co-Evolution: Closed-Loop Cognitive Mapping for Spatial Reasoning
Abstract
Scene-level spatial reasoning requires multimodal large language models to integrate partial video observations into task-sufficient spatial states. Existing cognitive-map pipelines support this integration but often lack reasoning-to-state feedback: externally supplied entity inventories leave relevance assessment outside the reasoning process, while fixed state estimates prevent evidence-driven revision when inference reveals uncertainty. To close this loop, we propose Reasoning-in-the-Loop Cognitive Mapping (RiCM), a framework in which reasoning directs state expansion and verification, and updated states guide subsequent inference. A dedicated spatial tool uses instance-consistent 3D anchors to ground entity association and evidence-gated local revision, decoupling geometric estimation from high-level reasoning. Final-answer rewards alone, however, do not explicitly supervise the sufficiency of intermediate states. We therefore introduce Progressive State Sufficiency Training (PSST), which promotes early and sustained state sufficiency through evaluation by an independent reader and unbiased estimation of the fraction of a fixed interaction horizon covered by sufficient states. Its reward combines joint correctness of the policy and the terminal-state reader with an interaction penalty, enabling the policy to learn when to expand, verify, and stop without ground-truth maps or oracle entity inventories. Experiments on ReVSI and MMSI-Video-Bench show consistent gains, supporting the value of bidirectional reasoning-state feedback for spatial reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.