acceptodds
Under review as a conference paper at ICLR 2027

Gaze World Model

Abstract

Human gaze provides a social cue between people and their surroundings, revealing where their visual attention is directed. Modeling this cue over time requires understanding where a person is looking in the observed scene and anticipating how their gaze will change as the scene unfolds. World models connect an understanding of the current environment with predictions of how it will evolve. Inspired by this perspective, we introduce the Gaze World Model (GWM), which unifies observed gaze understanding and future forecasting through latent scene prediction. Given a video clip and the person’s head box, GWM uses V-JEPA 2.1 and a temporal trunk to encode scene context and person-specific cues. From this representation, it estimates gaze locations and detects gaze shifts in the observed clip. It further forecasts future gaze locations and shifts using latent scene features predicted by the world predictor. To evaluate both current gaze understanding and future forecasting, we augment VideoAttentionTarget and ChildPlay with gaze-shift annotations, complementing spatial localization with event-level evaluation of person-involving target transitions. GWM achieves state-of-the-art performance on observed gaze-target estimation and shift detection on the evaluated benchmarks, and outperforms the compared baselines on future gaze localization and shift forecasting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.