CoordMotion: Learning Predictive Eye-Body Coordination States for Human Motion Forecasting
Abstract
Human motion forecasting is an important topic in the areas of computer vision, robotics, and human-aware artificial intelligence. Prior human behavioral studies have shown that eye–body coordination provides informative structure for understanding actions, intentions, and task progression. However, this relationship has not been explicitly explored for human motion forecasting. We present , a novel method that learns eye–body coordination as structured predictive behavioral states to forecast human motion. Specifically, (1) we learn structured eye–body coordination states that explicitly organize task-dependent eye–body action transitions from complementary semantic text and physical value; (2) we forecast the evolution of the learned coordination states and use the predicted future states to guide deterministic base prediction and horizon-dependent residual diffusion; and (3) we introduce a causal Coordination-Guided Verifier that leverages predicted coordination states to rank and select motion candidates, enabling coordination-aware test-time selection from multiple stochastic samples. We conduct extensive experiments on both deterministic and stochastic human motion forecasting tasks using the EE4D-Motion and Nymeria datasets. Experimental results show that our method outperforms state-of-the-art deterministic forecasting methods, reducing the mean per-joint position error by up to 9.7%, and produces higher-quality stochastic futures, reducing the average displacement error on EE4D-Motion by 20.2%. These results validate the significant potential of structured eye–body coordination states for human motion forecasting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.