DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation
Abstract
Physics-based character control for dexterous Human-Object Interaction (HOI) relies heavily on motion-capture data, which limits its applicability beyond captured scenarios. Meanwhile, video generative models can synthesize realistic HOI videos for unseen objects and scenarios directly from text prompts, but 3D interactions learned from such synthetic data remain kinematic without ensuring physical plausibility. Imitating these videos with physics-based control is a promising direction, yet unreliable object tracking and hand-object misalignment make it difficult to obtain accurate imitation targets. To address these challenges, we present DeVI (Dexterous Video Imitation), a novel framework that learns physics-based character control for dexterous HOI by imitating text-conditioned synthetic videos. To bypass unreliable 3D object motion reconstruction, we introduce hybrid imitation targets that combine reconstructed 3D human motion with tracked 2D object trajectories. Our Visual HOI Alignment further refines the human reference to align with both the generated video and the initial 3D object configuration, making it suitable for physical interaction. Extensive experiments show that DeVI's hybrid representation is more effective than those used by the baselines and that Visual HOI Alignment improves reconstruction quality and imitation success. We further demonstrate physically plausible functional manipulation from text instructions, including target-aware interactions and articulated object manipulation, without requiring motion-capture demonstrations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.