acceptodds
Under review as a conference paper at ICLR 2027

UniHOI: Video and Motion Co-Generation via a Unified Latent Space for Hand–Object Interaction

Abstract

Hand-object interaction (HOI) generation aims to synthesize physically plausible videos and 3D motion. Motion generation models capture structured 3D interaction dynamics but are limited by scarce high-quality data, while video foundation models learn rich visual priors yet lack explicit modeling of 3D geometric structure. To leverage these complementary strengths, recent works have explored video and motion co-generation frameworks that jointly model visual priors and 3D structural constraints. They often employ direct attention-based interactions to capture cross-modal dependencies between video and motion representations. However, these methods overlook the spatiotemporal granularity mismatch between heterogeneous representations, hindering fine-grained cross-modal correspondence and degrading native representations. Therefore, our key insight is to introduce a modality-shared yet interaction-comprehensive latent representation for cross-modal interaction. Motivated by this, we propose UniHOI, a unified video and motion co-generation framework that preserves modality-specific generation while constructing a unified latent space to abstract away modality-specific organization and retain complementary interaction information. Specifically, we organize this space along two complementary dimensions, visual dynamics and 3D structure, thereby establishing fine-grained spatiotemporal correspondence and enabling collaborative modeling of rich visual priors and structured 3D information. Furthermore, we introduce cross-modal rhythm alignment to promote consistent temporal dynamics between video and motion generation. Extensive experiments validate our superiority over state-of-the-art.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.