acceptodds
Under review as a conference paper at ICLR 2027

SCOPE: Learning Spherical Coordinate Grounding for Spatial Understanding in MLLMs

Abstract

Spatial understanding is pivotal for machines to advance toward physical world intelligence, yet current MLLMs still struggle to estimate object distances, judge relative directions, and reason about 3D layout. One line of work adds 3D encoders, which changes the architecture and weakens general multimodal ability. The other fine-tunes on spatial reasoning data, where supervision targets final answers rather than the 3D locations of relevant objects. Models may then rely on language priors without learning where objects are in 3D space. We therefore seek to train MLLMs to localize objects in 3D, which raises a question: which location representation best supports MLLMs' spatial reasoning? Under a protocol that varies only this choice across three MLLMs, we compare four representations common in 3D vision and find spherical coordinates the most effective across all three models. In spherical coordinates, distance and direction occupy separate axes: an observer's yaw rotation changes only azimuth, while moving directly toward an object changes only distance. To train MLLMs with spherical coordinate grounding, we develop an automatic pipeline that generates grounding samples from annotated 3D scans, yielding SphereQA, a large dataset of 300K samples in egocentric and allocentric reference frames. We then introduce SCOPE, an MLLM trained on SphereQA to predict spherical coordinates through per-axis classification with a spherical coordinate tokenizer instead of raw-text generation. Its two variants improve over their respective base MLLMs by 10.2/11.1 points on average across four spatial benchmarks, achieve the best average accuracy among the evaluated open-source and specialist MLLMs, and preserve general multimodal ability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.