An Anatomy-Informed 20D Representation for Text-Conditioned 3D Hand Pose Generation
Abstract
3D hand pose is widely used in sign language, human-computer interaction, and virtual avatars. Its representation determines how finger articulation is measured, controlled, and generated. 3D hand poses are commonly represented using MANO, which drives hand meshes with a 45D vector of local 3D joint rotations. This representation uses more dimensions than the hand's anatomical degrees of freedom, and its components do not directly represent finger flexion and spread angles. In this paper, we introduce AHand20, an anatomy-informed 20D hand pose representation guided by joint structure and degrees of freedom, with angles defined according to how each joint moves. AHand20 supports conversion to and from MANO and enables pose augmentation by varying finger flexion and spread while maintaining anatomical validity. Using the augmented poses, we construct a dataset containing 15,400 MANO poses and 30,800 descriptions across 154 handshapes, with each description linked to multiple pose variations. Trained on this dataset, a text-conditioned generator predicts 20D hand poses from descriptions of unseen handshapes and achieves lower joint position and rotation errors than existing methods. AHand20 also covers the main joint motions and converts to and from MANO with low error. Beyond pose generation, the representation identifies implausible poses in existing datasets, extending its application to hand pose data filtering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.