Interaction-Centric Human-to-Dexterous Demonstration Capture System and Dataset
Abstract
Large-scale human manipulation demonstrations offer a low-cost source of diverse data for dexterous robot manipulation, but the substantial embodiment gap between humans and robots limits their direct use for robot learning. We present an interaction-centric hand-object data capture pipeline that converts human demonstrations into dexterous manipulation data through cost-efficient acquisition and annotation. Our key insight is to leverage spatio-temporal hand-object interaction as an embodiment-transferable representation, using it to constrain motion estimation and preserve interaction relationships during human-to-robot transfer. Specifically, we introduce interaction-guided object motion estimation that exploits spatio-temporal contact plausibility to prune implausible pose hypotheses, complementing limited visual evidence under sparse-view capture and reducing the pose search space for efficient annotation. We further preserve hand-object contact relationships during simulation-based motion retargeting, transforming human demonstrations into physically plausible robot manipulation data. Experiments demonstrate several-fold improvements in object motion annotation efficiency over existing methods and sub-centimeter object pose estimation under sparse-view capture, while physics-based simulation validates interaction-preserving cross-embodiment transfer. Building on this pipeline, we construct H2D, comprising 5.6 hours of human demonstrations and 2,200 dexterous manipulation trajectories with structured 3D hand-object states.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.