Geo2Dex: Learning Multi-Embodiment Dexterous Retargeting from Human Videos via Feed-Forward 3D Reconstruction
Abstract
Learning from human manipulation videos offers a promising path toward scalable robot learning by reducing dependence on costly robot demonstrations. Accurate retargeting is central to this goal, as the morphological gap between human and robotic hands prevents direct reuse of demonstrated motions. In this paper, we present Geo2Dex, a novel two-stage framework that couples feed-forward 3D reconstruction with embodiment-conditioned trajectory prediction to learn retargeting from paired human–robot demonstrations. Geo2Dex retains reconstructed hand motion and latent scene geometry in a shared interaction representation, which guides a structure-conditioned slot decoder to generate wrist and joint trajectories across robotic hands. Comprehensive evaluations on HRDexDB and DexCap demonstrate marked improvements in reconstruction accuracy, retargeting performance, and simulation replay success. Using these trajectories for fine-tuning also substantially improves vision–language–action policy performance in simulation and on real robots.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.