acceptodds
Under review as a conference paper at ICLR 2027

HandWM: Hand–Object World Modeling with Explicit Interaction Dynamics

Abstract

Video world models aim to predict future visual observations conditioned on predefined controlling signals. While most existing methods focus on exploration (i.e., controlling the observation pose), the recent emergence of interactive world models begins to control the behavior of an agent in the video and model their interaction with the world. In this paper, we focus on hand–object interaction which poses a particular challenge because future observations depend on both 3D interaction dynamics and complex visual appearance. Direct action-conditioned video generation couples dynamics and appearance, leading to inaccurate interaction predictions. We propose a factorized interactive world modeling framework that disentangles explicit 3D interaction dynamics from visual observation synthesis. Given an initial observation and future hand controls, we first predict the 3D object geometry and the initial 3D hand pose and then leverage a simulator-based interaction module to produce an explicit 3D hand–object trajectory. We introduce pixel-aligned surface conditioning to rasterize future hand–object surfaces into the target view. The resulting maps provide pixel-aligned geometric guidance with explicit visibility and occlusion information. We further propose a bidirectional geometry–appearance interaction module to enable spatially aligned control and generation features to exchange information during synthesis. On DexYCB, our model reduces FVD from 31.9 to 16.9 and hand-pose MPJPE from 31.1 to 21.5 mm relative to the state of the art while preserving motion fidelity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.