GeoMotion-HOI: From 3D Interaction Motion to Controllable Hand–Object Videos
Abstract
Recent advances in hand–object interaction (HOI) video synthesis have improved visual realism, while 3D conditioning has strengthened control over motion and spatial structure. However, future 3D interactions are rarely available at inference time, and their motion, contact, and object-relative relations must remain coordinated to provide reliable video control. We propose GEOMOTION-HOI, a two-stage framework connecting future 3D interaction generation with controllable egocentric video synthesis. Conditioned on text, an observed interaction state, hand context, and object geometry, the first stage jointly predicts global motion, part-level contact, and object-relative relations. A dynamic binding module uses multiscale temporal convolutions to model evolving hand–object relationships, followed by contact–relation fusion and contact-aware residual refinement to produce a coordinated future 3D sequence. The second stage renders this sequence into two complementary conditions: Tracking Video conveys surface-point correspondences across frames, while Geometry Video combines depth-based surface structure and occlusion with stable, first-frame-derived colors that associate hand and object regions over time. A trainable control branch injects both conditions into a pre-trained image-to-video model, allowing the predicted sequence to guide motion and geometry while the observed first frame anchors scene appearance. Experiments on OakInk2 and TACO reduce motion FID by 40.9% and 22.7%, respectively, relative to the best evaluated baselines; video evaluations with given 3D conditions show improved fidelity, and ablations confirm that combining both conditions outperforms either alone across all five video metrics. GEOMOTION-HOI enables 3D-guided HOI video synthesis without an externally specified future trajectory or an exact target hand pose.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.