acceptodds
Under review as a conference paper at ICLR 2027

ContactFlow: A video action conditioning that transfers across embodiments

Abstract

World models offer a promising route toward more general embodied intelligence by enabling robotic agents to anticipate the consequences of actions before execution. Video Generation based world models are particularly attractive because they can learn rich spatiotemporal interaction dynamics from large-scale video data. By capturing aspects of real-world physics, these dynamics could allow the models to simulate how scenes evolve under different actions. Realizing this potential requires models to learn from interaction data at sufficient scale and diversity, most of which depict humans, and transfer the acquired knowledge to different robotic embodiments. However, such a transfer requires an action representation that provides a common interface across humans and robots. Existing approaches instead condition on embodiment-specific actions, poses, or kinematics, tying the learned dynamics to a particular actor or control space. This coupling limits both the use of heterogeneous interaction data and transfer across embodiments. We propose Contact Flow, an embodiment-agnostic action representation that encodes manipulation through the spatiotemporal trajectories of 3D contact points between an actor and a target object. Contact Flow captures the object-side interaction geometry that determines how actions affect the environment, providing a common interface across embodiments. We condition a video world model on Contact Flow to predict future scene evolution from an initial observation. Experiments across human and robotic manipulation benchmarks show that Contact Flow enables accurate video prediction and effective cross-embodiment transfer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.