RoboApproachBench: Diagnosing Spatial Understanding during Real-Robot Approach
Abstract
Multimodal large language models (MLLMs) are increasingly explored as perception and reasoning modules for embodied agents. Physical approach provides a concrete setting for evaluating their spatial understanding: as a robot moves toward a persistent target, target-relative geometry, ego-motion, and visibility change together. We introduce RoboApproachBench, a real-robot benchmark for diagnosing visual spatial understanding during recorded approaches. The benchmark contains 9,491 questions across 14 tasks, constructed from a corpus of 2,650 real-robot approach trajectories. Its tasks cover goal grounding, target geometry and approach progress, ego-motion, turn-sequence structure, and target visibility and bearing. Models receive task-specific monocular RGB observations and explicitly stated task information, while synchronized sensor measurements and target annotations support reference-answer construction. We evaluate 28 open- and closed-source MLLMs and observe uneven task-level performance. Across 28 MLLMs, peak-turn characterization exceeds temporal localization in 27 of the 28 models, while leading models’ strengths in approach ordering and visibility-event recognition coexist with weaker target-relative geometry. Across same-scale Qwen3-VL pairs, Thinking variants improve aggregate scores but exhibit task-specific regressions. Controlled bearing diagnostics on Qwen3.8-Flash yield higher accuracy with post-reappearance than pre-reappearance prefixes, highlighting the importance of observation endpoints when estimating earlier invisible states. RoboApproachBench provides a trajectory-grounded testbed for diagnosing task-specific strengths and limitations in MLLMs’ visual spatial understanding during real-robot approaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.