acceptodds
Under review as a conference paper at ICLR 2027

PLAN-Bench: Can Existing VLMs Go Well Beyond Perception to Reason and Plan in Embodied Worlds?

Abstract

Vision-Language Models (VLMs) have achieved strong performance in visual perception and spatial reasoning, leading to growing interest in using them as high-level planners for embodied AI.However, physical tasks require models to integrate spatial information, reason over temporal states, and generate multi-step plans under various constraints, raising the question of whether current VLMs can reliably support such reasoning and planning independently of execution.We introduce Plan-Bench, a benchmark for evaluating the high-level reasoning and multi-step planning capabilities of VLMs relevant to embodied worlds, while isolating the influence of action execution.It covers three complementary dimensions, spanning 16 subclasses and 699 curated samples.Extensive evaluation of 28 VLMs reveals room for improvement, with even the strongest model achieving only 75.0% accuracy.We further find that model scaling generally improves these capabilities.Ablation studies also indicate larger models show increasing potential as high-level planners.Overall, Plan-Bench provides a lightweight probe for evaluating VLMs' reasoning and planning capabilities relevant to embodied environments and assessing their potential as decision makers for physical tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.