Intrinsic Alignment by Generation: Monocular Hand–Object Reconstruction via Interaction Scaffolds
Abstract
Monocular hand–object reconstruction ultimately requires separate hand and object assets, but their relative pose and scale are inherently joint geometric quantities. This suggests the interaction is most naturally established in a shared geometry. Existing methods either incorporate hand geometry or interaction constraints into learned reconstruction while making limited use of object-centric image-to-3D priors, or exploit pretrained image-to-3D priors while deferring hand–object placement to test-time alignment optimization. JIGAR (Joint Interaction Generation and Asset Recovery) bridges this gap: its Joint Interaction Generation (JIG) stage adapts a pretrained image-to-3D model to jointly generate hand–object geometry as an interaction scaffold, which then guides separate asset recovery through in-place object regeneration and parametric hand recovery. This preserves the hand–object configuration while allowing the prior to refine object geometry. In this way, JIGAR combines pretrained generative shape priors with intrinsic hand-relative reconstruction without requiring test-time optimization of the hand–object relative placement. Evaluations on SHOWMe, HOT3D, and HO-Cap show that JIGAR outperforms all evaluated baselines in object and interaction reconstruction, reducing hand-aligned object Chamfer distance by 42.3–63.2% relative to the strongest evaluated baseline across the three datasets. The code is included in the supplementary material and will be open-sourced upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.