acceptodds
Under review as a conference paper at ICLR 2027

Generative Hand Modeling: A Unified Model for Egocentric Hand Motion Reconstruction, Infilling and Generation

Abstract

Modeling 4D hand motion is important for understanding human manipulation and learning from human demonstrations. Existing approaches, however, typically treat reconstruction, motion infilling, and generation as separate problems, despite their shared need to model hand articulation and temporal dynamics. We present Generative Hand Modeling (GHM), a unified framework that formulates these tasks as conditional hand motion modeling under varying levels of visual observability. GHM uses a shared motion representation and backbone to learn from different forms of visual and semantic supervision. It jointly models both hands within a common motion space. Experiments across these settings demonstrate strong reconstruction and temporal infilling performance, while retaining the ability to synthesize hand motion from scene and language conditions. Our results show the potential of unified generative modeling for learning hand motion from diverse sources of egocentric interaction data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.