Predictive Scanning for Efficient Spatiotemporal Representation Learning
Abstract
State-space models (SSMs) provide an efficient alternative for video modeling, but most existing approaches serialize visual tokens using fixed, content-independent scanning orders. Such predefined traversals are poorly matched to videos, where informative evidence is spatially sparse, temporally varying, and often highly redundant across frames. We propose **VScan**, an information-guided adaptive scanning framework that jointly determines *where to look, at what granularity, and in what order* visual regions should be processed. VScan consists of three complementary components: an information-aware scoring mechanism that combines static structure, residual motion, and motion-compensated predictive innovation; an adaptive spatial-resolution strategy that allocates finer representations to more informative content; and an information-guided spatiotemporal path planner that organizes selected regions into coherent and non-redundant trajectories. Rather than modifying the underlying state-space operator, VScan adapts the sequence presented to it according to the spatial and temporal information distribution of each video. Experiments on action recognition, fine-grained video understanding, and video semantic segmentation show consistent improvements, achieving **88.5%** Top-1 accuracy on Kinetics-400, **97.6%** accuracy on Breakfast, and **65.8%** mIoU on VSPW. Extensive ablations further demonstrate the complementary roles of information scoring, adaptive resolution, and path planning, while controlled scanning-policy comparisons show that VScan achieves a better trade-off between information acquisition, traversal continuity, and redundant computation than conventional scanning strategies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.