acceptodds
Under review as a conference paper at ICLR 2027

Moving Forward with Video Saliency: A New Dataset and Benchmark where Motion Matters

Abstract

Video saliency prediction attempts to capture more natural human visual behaviour than static image saliency, and is inherently harder to model due to the additional temporal dimension. Video saliency benchmarks rest on an implicit premise that predicting gaze on video requires utilizing temporal activity distributed across frames. Prior work has challenged this premise, showing that a deliberately static baseline recovers a significant fraction of the explainable gaze information on LEDOV without integrating temporal context, and that video saliency models fail in the same places as the static baseline. We first verify that this diagnosis still stands: under a more capable gold standard than the original analysis used, and an updated panel of recent architectures, the strongest temporal architecture in the panel still does not seem to substantially improve over a fine-tuned static baseline. However, it remains unclear whether the marginal gap reflects limitations of current temporal architectures or a lack of genuinely temporal patterns in the benchmark itself. We introduce SalTempto, a new training dataset and evaluation benchmark with greater dynamism: 224 one-minute action-centric clips drawn from HACS-Segments, with gaze recordings from up to 16 subjects. On SalTempto, the same fine-tuned static baseline recovers only about 13% of the available gaze information, which is a much smaller fraction than can be recovered on LEDOV. The strongest fine-tuned temporal architecture shows a substantial gain in performance over the static baseline, indicating that it does capture meaningfully more temporal information, which LEDOV fails to measure. Yet, even this SoTA model still leaves nearly half of SalTempto’s headroom unexplained, indicating that there is still a lot of room for improvement in the field of video saliency. In addition, manual examination of SalTempto recordings allows us to describe three new human visual tendencies: object permanence despite occlusion, scene inertia, and anticipatory saccades. SalTempto releases with raw visual recordings and training and evaluation videos. It will be released as a hosted video saliency benchmark based on the provided test-set videos upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.