Geometry-Guided Cross-View Reasoning in Multimodal Large Language Models
Abstract
Multi-view spatial reasoning requires MLLMs to associate observations of the same physical scene region across viewpoints. Yet existing geometry-aware MLLMs mainly enrich individual visual tokens with 3D or camera information, leaving such cross-view correspondences to emerge implicitly from appearance-driven attention. We introduce GeoMLLM, a geometry-guided MLLM that grounds cross-view information exchange in camera geometry. GeoMLLM equips each visual patch with a camera-ray representation and uses pairwise ray and epipolar cues to constrain which tokens can plausibly correspond across views. During training, depth-derived soft correspondences further supervise cross-view attention, teaching the model to route information between observations of the same physical scene region. Built on Qwen3-VL-4B, GeoMLLM achieves state-of-the-art performance on SPAR-Bench, reaching overall, with on multi-view tasks ( points) and on Cross-view Perception ( points). GeoMLLM also consistently improves over its matched backbone across additional multi-image and 3D spatial reasoning benchmarks. These results demonstrate that explicitly modeling camera geometry in cross-view token interactions provides a strong inductive bias for multi-view spatial reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.