StateVid: Video Representation as an Evolving Visual State
Abstract
Video encoders usually represent a video as a sequence of spatial features produced from individual frames. This frame-based representation treats each observation as a separate feature set, leaving the integration of persistent and changing visual content to later processing. We view frames as discrete observations of an evolving visual process, and its evolving state as the object of video representation. We propose StateVid, which updates this state causally at each spatial location with learned retention and writing gates. The state itself carries the video representation, and ordered snapshots expose its evolution to downstream models. The resulting visual state keeps a fixed size while accumulating evidence over time. StateVid learns these transitions from short video-text pairs through supervision of historical retention and adaptation to new content. StateVid demonstrates strong performance on video tasks compared to models trained at larger scales, while maintaining image capabilities. At the same time, it achieves better token efficiency and performance on VideoLLM. We additionally find that this representation method has stable extrapolation performance and exhibits selective attention to video content.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.