VA-Bench: From Spatial Reasoning to Embodied Action in General-Purpose VLMs
Abstract
Spatial intelligence in vision-language models is commonly evaluated through question answering over pre-specified visual observations. Robot manipulation offers a more demanding test, but existing evaluations typically assess learned visuomotor policies rather than the spatial intelligence of general-purpose VLMs. We introduce VA-Bench, which uses closed-loop embodied manipulation to evaluate whether general VLMs can turn spatial understanding into physical action. Given a task goal and visual demonstrations of the task procedure, the model must infer action-relevant geometry and generate metric robot actions without relying on a task-specific action policy or privileged spatial information. During execution, it can perform active perception by choosing additional viewpoints when needed and revise subsequent actions from interaction feedback. We refer to this overall capability as self-directed metric spatial grounding. VA-Bench further tests geometric shifts by preserving the task procedure while changing the action-relevant geometry. Across 14 tasks, existing VLMs show near-perfect target identification and task understanding, yet remain substantially weaker at spatial grounding and online correction; the highest overall task success rate is only 53.9%. Active perception improves success by 7.86–23.93% over controls given predefined views, while geometric shifts cause drops of up to 32.14%. Overall, these results show that for general VLMs, knowing what to do does not yet translate into knowing where and how to act reliably.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.