Grasping the Future: Joint 3D Hand–Object Forecasting from Internet Videos
Abstract
To interact with an object, we need to anticipate both how it will move and how our hands should move with it. Human videos capture these relationships across everyday activities, offering a natural source for learning to predict future interactions. Learning from them requires recovering motion from completed interactions and using it to anticipate what could happen next. We present Grasping-the-Future, a unified generative model of joint 3D hand-object motion, including wrist motion, finger articulation, and object position and orientation. Through masked spatiotemporal completion and flow matching, we train a single model to forecast future interactions, predict hand motion given object motion or vice versa, generate both trajectories toward a desired final object pose, and complete partially observed interactions. We explicitly model hand-object relations and incorporate contact constraints to help preserve grasp relationships during motion. To obtain 3D supervision from internet videos, we develop a scalable pipeline that combines pretrained perception models with joint geometric optimization to reconstruct completed interactions. We combine these recovered trajectories with annotated interaction datasets to learn future motion from partial observations. We demonstrate strong performance across all tasks, outperforming recent state-of-the-art models. Moreover, we show geometric and physical plausibility of the predicted trajectories by retargeting them to a robotic embodiment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.