GroundHand: Rethinking Egocentric Manipulation Data Collection with High-Precision Hand Annotation
Abstract
Egocentric dexterous manipulation data are increasingly used for robot manipulation learning, yet existing data collection pipelines often cannot provide high-precision hand state annotations with consistent absolute physical scale. Common approaches first estimate hand poses from egocentric videos, recover ego-camera poses separately, and then transform the estimated hand states into a world coordinate system. Errors can accumulate throughout this pipeline, rendering the resulting annotations unreliable for fine-grained, dexterous manipulation. In this work, we propose GroundHand, a low-cost capture-and-annotation system that captures synchronized ego-view and wrist-view manipulation videos and produces high-precision MANO hand states directly in a unified workspace coordinate system. GroundHand uses a calibrated multi-view RGB-D setup with an exo-view annotation hardware cost of approximately $3,000, substantially lower than Vicon-based MoCap systems. Built upon this system, we develop a multi-source hand annotation model to estimate fine-grained hand pose and global translation. To continuously improve automatic annotation quality, we further introduce iterative human-in-the-loop model refinement, where sparse manual corrections are fed back during training to progressively refine the model. Experiments on existing benchmarks and our manually validated data show that GroundHand improves hand state accuracy and annotation efficiency over baseline methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.