Gaze World Model
Abstract
Human gaze provides a social cue between people and their surroundings, revealing where their visual attention is directed. Modeling this cue over time requires understanding where a person is looking in the observed scene and anticipating how their gaze will change as the scene unfolds. World models connect an understanding of the current environment with predictions of how it will evolve. Inspired by this perspective, we introduce the Gaze World Model (GWM), which unifies observed gaze understanding and future forecasting through latent scene prediction. Given a video clip and the person’s head box, GWM uses V-JEPA 2.1 and a temporal trunk to encode scene context and person-specific cues. From this representation, it estimates gaze locations and detects gaze shifts in the observed clip. It further forecasts future gaze locations and shifts using latent scene features predicted by the world predictor. To evaluate both current gaze understanding and future forecasting, we augment VideoAttentionTarget and ChildPlay with gaze-shift annotations, complementing spatial localization with event-level evaluation of person-involving target transitions. GWM achieves state-of-the-art performance on observed gaze-target estimation and shift detection on the evaluated benchmarks, and outperforms the compared baselines on future gaze localization and shift forecasting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.