Geo-REPA: Geometry-Aware Representation Alignment Enhances Native Spatial Intelligence in MLLMs
Abstract
Multi-modal Large Language Models (MLLMs) excel at semantic understanding but often struggle to reason about physical size, depth, and spatial relationships in 3D environments. Visual geometry encoders provide strong geometric priors, yet our analysis of frozen checkpoints on held-out samples finds that correspondence between MLLM hidden states and geometry features weakens in later layers after input-level fusion, which lacks direct intermediate supervision. We introduce Geo-REPA, a geometry-aware representation alignment framework that encourages intermediate MLLM features to preserve and use 3D cues. To account for camera motion between image and 3D scene coordinate frames, Geo-REPA introduces a pose-aware SE(3)-equivariant inductive bias. Learned pose tokens predict per-frame rotations and translations, which induce a transformation of intermediate visual tokens. A dual alignment objective supervises the transformed visual tokens with 3D geometry features and the pose tokens with camera pose targets from a visual geometry encoder. Across VSI-Bench, SQA3D, and Scan2Cap, Geo-REPA achieves competitive or leading spatial performance, with Scan2Cap gains extending to a newer MLLM backbone. Representation diagnostics associate stronger alignment with lower depth error, while cross-view visualizations show reduced separation among views after transformation. Quantitative analyses associate alignment with spatial performance, show pose sensitivity and meaningful rotation and translation direction estimates. Matched ablations favor pose-aware alignment over no alignment and direct matching across three tasks. On general multimodal benchmarks, the evaluated variant scores modestly below its original backbone. Neither the external geometry encoder nor the alignment modules are needed at inference. Code, scripts, and checkpoints will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.