acceptodds
Under review as a conference paper at ICLR 2027

HandCrafter: Monocular 3D Hand Motion Reconstruction with Multi-View Video Diffusion Priors

Abstract

Monocular 3D hand motion reconstruction remains challenging in hand-object interaction videos due to persistent occlusions and the resulting spatial ambiguity. We present HandCrafter, a generative multi-view framework that reconstructs 3D hand motion by synthesizing complementary viewpoints from a single monocular video. We first leverage an off-the-shelf geometry foundation model to warp the input video into novel viewpoints, producing incomplete but spatially aligned conditions that are subsequently completed by an adapted video diffusion model. Hand motion is then recovered from the resulting multi-view diffusion representations. Instead of relying on hand detection and cropping, our estimator uses sparse 2D joint anchors for spatial localization and decodes unified MANO parameters from cross-view diffusion representations, enabling crop-free 3D hand motion reconstruction. To mitigate the scarcity of multi-view HOI data, we further mix limited multi-view data with scalable monocular supervision, enabling cross-view and temporal modeling to benefit from a broader training distribution. Extensive experiments demonstrate superior performance under severe occlusions and strong generalization across diverse in-the-wild HOI scenarios.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.