SYMBOLIC COMPOSITIONAL PLANNING FROM DEMONSTRATED EXPERIENCE FOR EMBODIED AGENTS
Abstract
Vision-language models (VLMs) exhibit strong in-context learning capabilities, yet embodied agents built on them struggle with complex tasks given only a few demonstrations, often omitting critical steps, violating action preconditions, or organizing the overall workflow incorrectly. These failures stem from an incomplete understanding of demonstrated action dependencies and a failure to distinguish reusable local procedures from task-specific workflow structure. To address this problem, we propose an embodied planning method that combines trajectory composition with symbolic analysis, reusing demonstrated local procedures while reorganizing them into workflows for new tasks. Demonstration fragments preserve local action structure, while symbolic analysis makes their preconditions and effects explicit to guide composition and connections between fragments. During execution, the agent maintains scene memory and uses execution feedback to revise the remaining plan. Experiments on EB-ALFRED and EB-Habitat show success-rate improvements of 21% and 22 %, respectively, over standard few-shot planning with the same backbone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.