acceptodds
Under review as a conference paper at ICLR 2027

Zeva-EgoAction: Scalable Egocentric Mid-Training for Robot Manipulation

Abstract

Egocentric human videos contain manipulation experience, but their lack of action annotations limits their use in robot policy training. We propose Zeva-EgoAction, which extracts supervision from action-unlabeled egocentric videos to enhance robot-pretrained vision-language-action (VLA) models. Its Action-Centric Encoder (ACE) uses annotated trajectories to learn a physically grounded latent space through environment distillation and action-neighborhood matching, then generates action pseudo-labels for unseen and unlabeled Ego data. On an unseen dataset, ACE labels all evaluated samples without fine-tuning and improves action-representation fidelity, measured by action-distance correlation, by 37.6% relative to HaMeR with retargeting on the common valid subset. Scaling Ego mid-training data to 10,000 hours raises RoboTwin success from 63.8% to 75.3%; under this setting, approximately 4–5 hours of Ego video yield training gains comparable to one hour of robot demonstrations. Real-robot post-training further shows that task-matched Ego videos improve average success by 35 percentage points when Robot demonstrations are scarce. These results demonstrate that physically grounded pseudo-labels can convert large-scale human videos into effective supervision for robot policy training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.