Handmimic: Learning Dexterous Manipulations from Videos
Abstract
While humans naturally learn dexterous manipulation by watching others and practicing, robotic teleoperation of dexterous hands remains prohibitively expensive and non-trivial. We introduce HandMimic, a framework for learning dexterous manipulation policies directly from real-world monocular videos. We first reconstruct and align the observed hand-object interaction in simulation, while treating the resulting motion as an imperfect reference rather than a physically valid robot demonstration. We then practice in simulation using reinforcement learning to recover physically feasible behaviors from the reconstructed reference. To reduce dependence on a single reconstructed object, we further train across digital siblings with varying geometry and scale. Quantitative experiments show that policies learned from reconstructed references approach those trained from ground-truth references, while digital-sibling training improves generalization to unseen geometry and scale variations. We further validate the resulting motions on a real dexterous robot. Together, these results support using reconstructed human interactions as references for physical practice across varying object geometries.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.