Zeva-EgoAction: Scalable Egocentric Mid-Training for Robot Manipulation
Abstract
Egocentric human videos contain manipulation experience, but their lack of action annotations limits their use in robot policy training. We propose Zeva-EgoAction, which extracts supervision from action-unlabeled egocentric videos to enhance robot-pretrained vision-language-action (VLA) models. Its Action-Centric Encoder (ACE) uses annotated trajectories to learn a physically grounded latent space through environment distillation and action-neighborhood matching, then generates action pseudo-labels for unseen and unlabeled Ego data. On an unseen dataset, ACE labels all evaluated samples without fine-tuning and improves action-representation fidelity, measured by action-distance correlation, by 37.6% relative to HaMeR with retargeting on the common valid subset. Scaling Ego mid-training data to 10,000 hours raises RoboTwin success from 63.8% to 75.3%; under this setting, approximately 4–5 hours of Ego video yield training gains comparable to one hour of robot demonstrations. Real-robot post-training further shows that task-matched Ego videos improve average success by 35 percentage points when Robot demonstrations are scarce. These results demonstrate that physically grounded pseudo-labels can convert large-scale human videos into effective supervision for robot policy training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.