acceptodds
Under review as a conference paper at ICLR 2027

Foresight to Motion: Goal-Conditioned Autoregressive Human Motion Prediction from Egocentric Video

Abstract

Generating plausible human motion from egocentric observations is important for human behavior understanding and embodied robotics. However, in first-person settings, most of the body remains outside the camera's field of view, making motion prediction from partial visual observations and egocentric camera motion challenging. Existing approaches rely on RGB-conditioned diffusion models, which are limited by sparse behavioral cues and fixed-horizon generation. To address these limitations, we introduce Foresight to Motion (FOM), a multimodal autoregressive framework for goal-driven, human motion prediction from egocentric video. Beyond visual observations, our model incorporates geometric and camera-motion cues together with explicit future interaction goals, providing guidance toward desired targets over extended temporal horizons. We further propose a three-stage training paradigm based on teacher forcing, self-forcing, and head-camera consistency, improving closed-loop robustness and reducing the train–inference discrepancy. Extensive experiments demonstrate strong performance on egocentric motion prediction benchmarks and improved cross-dataset generalization. Our goal-driven formulation also provides a flexible foundation for downstream human behavior modeling and video-conditioned motion planning for embodied systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.