DepthThinker: Towards Real-World Spatial Perception with Metric Depth Reasoning
Abstract
Accurate depth perception is essential for vision-language models (VLMs) in navigation, embodied interaction, and autonomous driving, where decisions depend on precise distances and spatial constraints. However, existing depth training methods largely rely on regression losses or token-level cross-entropy, providing limited supervision for how depth estimates should be used in spatial reasoning. As a result, fluent explanations may still be inconsistent with scene geometry and physical constraints, limiting their reliability in unfamiliar environments. We introduce DepthThinker, a VLM designed to support test-time scaling in spatial perception through extended, depth-aware reasoning. Following joint supervised fine-tuning on depth and reasoning tasks, we introduce D-GRPO, which uses verifiable task rewards and feedback on reasoning quality to improve the accuracy and consistency of spatial judgments. To evaluate how depth perception supports practical spatial understanding and decision-making, we introduce WorldDepthBench, a benchmark of 3,360 question-answers covering depth perception, geometric reasoning, and action decisions, alongside a general spatial understanding category. With only 2B parameters, DepthThinker achieves a mean of 0.915 across seven datasets, surpassing reported results from depth-specialized and proprietary VLMs, and reaches 0.927 in a four-dataset comparison with specialist depth estimators. On WorldDepthBench, it achieves 45.77% accuracy, improving over joint supervised fine-tuning by 6.85 percentage points and outperforming Qwen3-VL-8B-Thinking by 3.30 points. These results show that combining depth supervision with reasoning and verifiable rewards improves spatial understanding and the use of depth information in practical decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.