EgoJEPA: Beyond Retargeting for Dexterous Manipulation from Egocentric Videos
Abstract
Egocentric human videos offer a scalable alternative to costly robot demonstrations, yet transferring their manipulation priors to robot control remains challenging. The commonly adopted retargeting-based approaches convert human wrist and finger poses into embodiment-specific action commands through hand-crafted geometric mappings. However, such mappings require meticulous recalibration across different embodiments and risk propagating hand-tracking errors into downstream robot control pipelines. To this end, we propose EgoJEPA, a three-stage framework that learns embodiment-agnostic representations from egocentric human videos without rule-based human-to-robot retargeting: 1) Human-action pre-training learns effect-aware latent actions through JEPA-style visual prediction conditioned on hand-motion differences; 2) Latent-action-based world modeling jointly predicts future visual dynamics and latent actions, which decouples human prior pre-training from robot-specific control parameterizations; 3) Embodiment-specific tuning adapts the pretrained model to diverse robots through lightweight action adapters. Experiments on five real-world manipulation tasks across three robot embodiments demonstrate the effectiveness and efficiency of EgoJEPA. It achieves a 74% average success rate under the out-of-domain evaluation, compared with 65% for the retargeting baseline, while reducing cumulative cross-embodiment wall-clock cost by 47.2%, from 743.2 to 392.7 hours.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.