acceptodds
Under review as a conference paper at ICLR 2027

Video Foundation Models Are Native Spatio-Temporal Field Forecasters

Abstract

Natural videos and urban spatio-temporal fields exhibit similar patterns of spatial evolution, suggesting that video pretraining may provide useful forecasting representations. However, the benefits of video pretraining and the feasibility of sharing a backbone across heterogeneous urban domains remain insufficiently explored. We introduce VFM4ST, a framework that renders numerical fields as pseudo-videos and jointly encodes spatial and temporal structure using a pretrained video foundation model. Domain-specific three-port decoders separate forecast baselines, field evolution, and output geometry. Direct access to native observations retains local information for forecasting on both dense grids and sparse stations. Across traffic flow, radar nowcasting, crime, and air quality, one shared 87M-parameter backbone achieves better mean performance than the strongest evaluated baseline in every domain. The shared model remains within 2.2% of independently specialised models in each domain. Under the same joint training protocol, video pretraining outperforms image pretraining and random initialisation across all four domains. Frozen-backbone and low-data evaluations further support the value of pretrained video representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.