OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment
Abstract
Multimodal large language models (MLLMs) still struggle with spatial reasoning that requires perspective transformation. Their predictions often favor camera-centric cues over the reference object's viewpoint, leading to systematic errors in non-camera reference settings. We analyze this behavior on binary left/right conflict samples and find that camera-centric answer preference varies with reference-object orientation. To address this issue, we propose OrientSAM, an orientation-aware spatial alignment framework for multimodal models. OrientSAM injects explicit orientation information into multimodal representations through orientation-aware tokens and Fourier-based angle encoding, and further adopts a curriculum learning strategy to progressively improve perspective-aware reasoning. In addition, we build a spatial data construction pipeline to generate orientation-aware spatial supervision from large-scale images. Experiments on Spatial-MM, ViewSpatial, and 3DSRBench show that OrientSAM delivers its strongest gains on non-camera-view, person-centric, and orientation-sensitive tasks. Paired conflict-subset comparisons show reduced camera-centric answer preference, supporting explicit orientation-aware spatial inputs for allocentric reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.