acceptodds
Under review as a conference paper at ICLR 2027

PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation

Abstract

World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate actions. Existing approaches typically represent the world as RGB frames or latent counterparts while predicting actions as end-effector poses or joint angles, but they often struggle to capture the 3D spatial structure and contact geometry central to dexterous manipulation. We introduce Point World Action Model (PointWAM), a 3D world action model that decomposes the world into a scene (*i.e.*, environment) and hands (*i.e.*, actor), and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame. This explicit, disentangled representation enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection. Given a colored point cloud and a language instruction, PointWAM predicts how the scene and hands co-evolve in 3D space over time, then retargets the forecast hand motion to robot actions. On ten DexJoCo tasks, PointWAM surpasses the prior state-of-the-art average multi-task success rate by 11.7 percentage points, with scene-trajectory supervision contributing 10.9 points and human-data pre-training 56.9 points. On a real robot, PointWAM also outperforms GR00T N1.6 and on both single-arm and bimanual tasks, by up to 25 percentage points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.