acceptodds
Under review as a conference paper at ICLR 2027

EgoHEP: Predictive Hand-Eye Pretraining for Robot Visual Representation Learning

Abstract

Robotic manipulation requires visual representations that support the selection of relevant objects and regions for interaction. A central challenge is to learn from the way humans actively select visual information and coordinate it with subsequent hand movements. We introduce EgoHEP, a predictive hand-eye pretraining framework that learns this relationship from gaze and hand-wrist motion in egocentric recordings. EgoHEP combines hand-wrist reconstruction and forecasting with spatial attention predicted under gaze supervision, then aligns attention-weighted visual features with representations of future hand-wrist trajectories across multiple horizons, linking visual selection to manipulation dynamics. Experiments in simulation and on real robots demonstrate the transferability of visual representations learned through predictive hand–eye supervision, with EgoHEP achieving state-of-the-art performance on multiple downstream tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.