Learning World Representations from Future Physical Outcomes
Abstract
Multi-stage manipulation requires early actions to leave suitable conditions for later contact. We study how to learn world-model states that retain physical information needed for subsequent actions under limited interaction. Small pose differences can change later outcomes, motivating supervision that describes action consequences. We train a shared state to predict outcome distributions under different candidate action sequences. A generative model jointly predicts physical outcomes and next states, with training through generated states for sequential prediction and feedback planning. Across three manipulation tasks, matched downstream models show that physical supervision improves held-out prediction and mean task success from frozen representations. The full system improves success over direct task-value prediction by 7.3 to 11.1 percentage points on jointly unseen objects and action combinations. Further tests demonstrate useful prediction from generated states, frozen dynamics reuse for new programs, and source-dependent representation transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.