EgoWeave: Structured Motion Supervision from Calibrated Egocentric Video
Abstract
Egocentric demonstrations could provide scalable human-motion supervision for robot learning, but camera projection and motion reference frames make recordings difficult to reuse across devices. We present EgoWeave, which converts heterogeneous egocentric video into structured motion supervision. A hand-centered virtual camera and crop-consistent decoding let one frozen hand estimator process monocular and stereo pinhole or fisheye video. Timestamped reference poses then place the recovered head, wrists, and hand landmarks in a common coordinate system. We organize this motion as head movement, future-head-relative wrists, and normalized wrist-local articulation, and train with a compositional landmark objective before robot-native finetuning. In controlled fixed-WiLoR experiments with annotated hand regions, hand-centered viewing reduces camera-space MPJPE by 52.9% on Aria and 29.8% on Quest 3 in the – off-axis range. In the evaluated short-budget setting, the complete EgoWeave recipe improves task success over direct finetuning by 5.20–20.50 percentage points across four robot suites; under structured supervision, replacing official annotations with EgoWeave's reconstructed motion further improves all four suites. Code is available at https://anonymous.4open.science/r/EgoWeave/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.