Self-Correcting Imaginative Planning for VLM Embodied Agents
Abstract
Recent advances in Vision-Language Models (VLMs) have demonstrated their remarkable capability to serve as embodied agents in interactive environments. However, these agents often make plans (i.e., action sequences) that appear linguistically plausible but fail upon execution, without considering the visual consequences of each action. While prior works have explored enhancing VLM agents through language supervision or self-correction with text-based verification, enabling them to self-correct invalid plans against visual consequences before execution remains an open challenge. To tackle this, we propose Self-Correcting Imaginative Planning (SCIP), a training framework that enables VLM agents to self-correct their plans against imagined visual consequences before execution. SCIP integrates an imagination-guided policy learning scheme that grounds the VLM agent's planning in visual consequences, and a verification-driven imagination learning scheme that calibrates the imagination process to faithfully reflect the actual environment dynamics. Extensive experiments show that SCIP outperforms existing embodied VLM agents across multiple embodied planning tasks. The code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.