acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Projective Guidance for Unified Video and 3D Hand Motion Generation

Abstract

Generating paired egocentric video and 3D hand motion can provide supervision for dexterous robot learning. Recent video–action models jointly generate both modalities, avoiding error propagation from sequential generation. However, they learn cross-modal correspondence implicitly, so the generated video and hand motion are often individually plausible yet spatio-temporally inconsistent. For example, given the same instruction, the 2D video hand trajectory may not align with the reprojected 3D hand motion trajectory on the image plane. To address this limitation, we introduce Uni-HaVi, a unified model for 3D hand motion and video generation. Compared with 3D hand motion generation, which gradually stabilizes in the later flow-matching steps, we observe that the hand trajectory in the generated video is largely determined within the first few steps. Exploiting this asymmetry, we propose Dynamic Projective Guidance (DPG), which uses a video hand locator and a learned adapter to progressively guide 3D hand motion generation toward the video hand trajectory throughout the flow-matching process. As a result, DPG enforces stronger spatio-temporal consistency between the video and 3D hand motion, while allowing their appearance and articulation to be independently refined during generative process. On EgoDex, DPG reduces wrist alignment error from 42.3 to 19.0 px and improves whole-hand IoU from 0.245 to 0.313, as measured by wrist and hand evaluators. By treating either modality as given or generated, Uni-HaVi supports video-to-hand-motion, hand-motion-to-video, and joint generation with a single model, and achieves strong video–hand spatio-temporal alignment and improved 3D hand-motion quality while maintaining competitive video quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.