EgoAhead: Egocentric Human Motion Understanding with Joint Video–Motion Prediction
Abstract
Understanding and anticipating human motion from egocentric observations is a key task for proactive AR/XR and embodied assistants. Existing egocentric behavior understanding models are typically trained only to describe the current motion, while methods that forecast future motion rely on either hints from the future recording or precomputed 3D environment scene features, limiting their applicability in real-world settings. To address this gap, we introduce EgoAhead, a unified multimodal model trained jointly for motion description and future video and pose prediction. Extending from an unified backbone Lance, EgoAhead allows multimodal generation of future video frames and poses to co-supervise a shared representation alongside language-based motion description during training. At inference, the model can output language descriptions without generating video or pose tokens, preserving inference efficiency. On the Nymeria benchmark, EgoAhead outperforms baseline models on both current-motion and future-motion descriptions despite using a smaller backbone. Our findings demonstrate that multimodal generation supervision provides an effective way to strengthen egocentric behavior understanding for wearable AI assistants.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.