StereoVLM: Binocular Vision-Language Models for Metric 3D Spatial Reasoning
Abstract
Vision-language models (VLMs) exhibit strong semantic understanding but remain limited in their ability to reason about metric 3D space in the physical world. Existing VLMs typically infer 3D scene structure from one or more images without known camera calibration or cross-view geometric constraints. This makes metric 3D inference inherently ambiguous and leaves it unclear whether their spatial predictions are grounded in visual evidence or driven primarily by semantic priors. In contrast, calibrated stereo provides explicit cross-view geometry that resolves this ambiguity, offering a more reliable foundation for metric 3D perception and geometric generalization. We investigate how VLMs can learn transferable 3D representations directly from calibrated binocular observations. To this end, we introduce **StereoVLM**, the first geometry-aware binocular VLM for metric 3D spatial reasoning. **StereoVLM** follows a hierarchical learning paradigm spanning from cross-view correspondence and disparity estimation to calibration-conditioned metric depth perception and semantic spatial reasoning. We further find that disparity supervision generalizes substantially better than direct binocular depth supervision on unseen environments. To evaluate these capabilities, we introduce **Stereo3D-Bench**, a comprehensive benchmark across three levels, from point-level geometry to object-level spatial reasoning. **StereoVLM-8B** achieves 95.12% depth accuracy () and a 66.98 average object-level score on Stereo3D-Bench. Binocular training yields average gains of 3.54 points across five spatial benchmarks and 4.3 percentage points in RoboCasa manipulation success. Intervention results further indicate that these gains depend on valid cross-view geometry and cannot be explained solely by additional visual tokens or semantic priors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.