Hand-: Scaling Training Data for 3D Hand Pose Estimation via Geometry-Preserving Image Generation
Abstract
The performance and generalization of monocular 3D hand pose estimation remain constrained by the limited scale and diversity of existing training data. Although graphics-based synthesis provides accurate geometric annotations, conventional rendering pipelines rely on limited HDR backgrounds and predominantly centered-hand configurations, restricting both visual and perspective diversity. We introduce Hand-, a large-scale synthetic hand dataset with substantial diversity in hand poses, backgrounds, hand–object interactions, hand locations, and camera configurations. We construct the dataset through a geometry-preserving rendering-and-generation pipeline. Starting from hand poses obtained from existing datasets and generative pose models, we render textured meshes under diverse virtual camera settings. Guided by structured prompts, a large image-generation model transforms these renderings into photorealistic images with diverse appearances while preserving the underlying hand geometry. We further preserve geometric fidelity through rendering-based appearance control, geometry-constrained prompting, and geometry-consistency filtering. The resulting dataset comprises approximately 10.5 million images spanning isolated hands, hand–object interactions, third-person views, and egocentric scenarios, with accurate MANO, keypoint, mesh, and camera annotations. We further develop WiLoR by introducing camera-conditioned inputs and full-image perspective projection with the actual camera intrinsics. Experiments show consistent performance gains across benchmarks as the scale of Hand- increases. Trained with the full Hand- dataset, WiLoR achieves state-of-the-art performance on both FreiHAND and HO-3D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.