acceptodds
Under review as a conference paper at ICLR 2027

Seeing Is Not Planning: When and How Egocentric Signals Help Human Motion Forecasting

Abstract

Egocentric signals such as eye gaze and scene geometry are increasingly fed to human motion forecasting models, yet the reported gains are small and inconsistent across benchmarks. We ask when and why these signals help. On a controlled testbed spanning human–object (MoGaze), human–scene (GIMO) and human–human (EgoBody) interaction with two state-of-the-art forecasters, we train each network with and without egocentric input and study the gain. We find three answers. When: the gain peaks at input and output windows that cover the lag between gaze and body motion, which differs across datasets; a poorly chosen window eliminates the gain or makes it negative. Where: the gain concentrates in specific motion states and joints and is diluted or cancelled by aggregate metrics; on a pick-and-place benchmark the gain at the acting wrist is nearly twice the reported aggregate. What: current datasets under-supply the information egocentric signals need; oracle probes that provide the missing information substantially increase the gain; on EgoBody, a simple wearer status oracle raises the Global MPJPE gain from 4% to 14%. But what is missing is highly dataset-specific. Together, the three findings reframe a near-zero aggregate gain not as evidence that egocentric signals are useless, but as a consequence of poorly matched windows, indiscriminate aggregation and incomplete data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.