IntentWAM: Task Intent as Context for Cross-Task Generalization in World-Action Models
Abstract
In-context learning from demonstrations offers a promising route to unseen robot tasks, but transferring demonstrated behavior across embodiments requires connecting task intent to the current execution scene. We propose IntentWAM, a world–action modeling framework with a unified conditioning interface that combines an Object-Centric Relational Task Specification (ORTS) with target-scene role bindings. ORTS encodes task roles, endpoint scene graphs, and object dynamics from both robot and human demonstrations. Combined with target-scene role bindings, ORTS provides intent tokens that guide visual-future prediction and subsequent action generation. Policy post-training associates demonstrations and robot executions by task identity, without requiring trajectory-specific demonstration construction or temporal correspondence between them; deployment uses external demonstrations as context without parameter updates. On seven held-out RoboTwin tasks, IntentWAM achieves mean success across robot and generated human demonstrations with a fixed UR5-WSG execution policy. With filmed human demonstrations from each evaluated layout, real-world evaluation after task-specific adaptation yields 47.50% success on held-out layouts, versus 22.50% for the strongest evaluated baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.