ReL-WM: Grounding World Model in Progress-Aware Scene Graphs
Abstract
Successful robot manipulation depends on how interactions between entities evolve toward task completion, yet standard learned world models do not explicitly represent this structure. We introduce ReL-WM (Relational Latent World Model), a framework that grounds latent dynamics in progress-aware scene graphs. These scene graphs retain task-relevant entities and describe their physical interactions, spatial configurations, and affordance compatibility, while temporal qualifiers capture intermediate relational changes. Using scene graph supervision, ReL-WM learns a semantic state that evolves alongside visual dynamics during imagination. Relational milestones extracted from successful demonstrations define a task-progress potential, whose predicted changes guide policy learning. Our experiments show improved success over visual and object-centric model-based baselines across tabletop and mobile manipulation, together with stronger transfer to new objects and zero-shot generalization to unseen layouts and lighting conditions. These findings support explicitly modeling interactions and task advancement as a useful foundation for world models in robot manipulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.