acceptodds
Under review as a conference paper at ICLR 2027

Geometry-Guided Cross-View Reasoning in Multimodal Large Language Models

Abstract

Multi-view spatial reasoning requires MLLMs to associate observations of the same physical scene region across viewpoints. Yet existing geometry-aware MLLMs mainly enrich individual visual tokens with 3D or camera information, leaving such cross-view correspondences to emerge implicitly from appearance-driven attention. We introduce GeoMLLM, a geometry-guided MLLM that grounds cross-view information exchange in camera geometry. GeoMLLM equips each visual patch with a camera-ray representation and uses pairwise ray and epipolar cues to constrain which tokens can plausibly correspond across views. During training, depth-derived soft correspondences further supervise cross-view attention, teaching the model to route information between observations of the same physical scene region. Built on Qwen3-VL-4B, GeoMLLM achieves state-of-the-art performance on SPAR-Bench, reaching overall, with on multi-view tasks ( points) and on Cross-view Perception ( points). GeoMLLM also consistently improves over its matched backbone across additional multi-image and 3D spatial reasoning benchmarks. These results demonstrate that explicitly modeling camera geometry in cross-view token interactions provides a strong inductive bias for multi-view spatial reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.