PointDex: Dexterous Grasp Generation via Point-based Hand-Object-Contact Modeling
Abstract
Dexterous grasp generation requires precise reasoning over the spatial and functional relationships among hands, objects, and their contacts. Yet existing approaches often struggle to model these interdependent entities effectively, and provide only limited means for steering the generated grasps. We introduce PointDex, an effective, steerable framework for dexterous grasp generation that captures hand–object interactions by representing grasp pose, object contact, and hand contact as a unified set of 3D points in a structured space. This explicit, disentangled yet jointly shared representation allows a single transformer to simultaneously denoise all point types, and naturally supports flexible steering by conditioning on any subset of points while denoising the rest. Given a 3D object and a language instruction, PointDex jointly generates posed hand keypoints and paired object–hand contact points, and subsequently recovers wrist pose and joint angles via inverse kinematics. On the DexGYS and Dexonomy benchmarks, PointDex improves language-conditioned grasp success by 16.6 points over the strongest prior method and reduces steering contact distance by up to 80%, while unifying diverse steering modes within a single model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.