Rethinking Spatial Reasoning in MLLMs: Learning to Form and Use Depth
Abstract
Spatial reasoning is essential for multimodal models to understand scenes and support embodied interaction, yet representing spatial information does not ensure that it guides their answers. We investigate how models can learn to form useful depth representations and turn them into evidence for spatial judgments. Our finding that depth information becomes readable before it has its strongest influence on answers motivates a training strategy that couples early depth formation with its subsequent use. Using existing question-answering annotations, we supervise object depth in early layers and train later layers to answer consistently with controlled changes in represented depth relations, requiring no additional geometry data, encoder, or teacher. Across VSI-Bench, CV-Bench, BLINK spatial tasks, MMSI-Bench, and ViewSpatial-Bench, our method raises mean scores by 3.8, 1.7, and 2.5 percentage points over supervised fine-tuning on identical data for Qwen2.5-VL-3B, Qwen3-VL-4B, and Qwen3.5-4B, respectively. These findings support early depth formation and downstream use as complementary training targets for improving spatial reasoning while preserving the original inference architecture.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.