Learning to Plan with Natural-Language World Models from Human Egocentric Videos
Abstract
Human demonstrations provide examples of how tasks are performed, but embodied agents need reusable general knowledge of when actions apply, how they change the world, and how they can be recomposed into new plans. We present Planning in Abstract Worlds (PAW), a framework that learns natural-language abstract world models from human egocentric videos and fine-tunes a vision-language model (VLM) to plan with them in a closed loop. PAW represents tasks with natural-language predicates and high-level actions with explicit preconditions, effects, and bimanual coordination types. Instead of treating the induced abstract world model as an exact transition system, it uses the model to reconstruct state-action supervision from unlabeled video and as structured context for a learned planner. The planner maintains a belief state, supports information gathering under partial observability, and replans from new visual observations. Across offline planning and online human assistance experiments in three everyday manipulation domains, PAW outperforms VLM planner baselines and symbolic planning with induced models, generalizes to unseen action compositions, and remains effective with incomplete models. Robot demonstrations further illustrate transfer from human videos to physical execution. Together, these results support learning reusable abstract world models from human videos and learning to plan with them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.