VLM4ST: Repurposing Vision-Language Models for Spatiotemporal Forecasting
Abstract
Spatiotemporal (ST) forecasting requires modelling spatial dependencies and temporal dynamics in high-dimensional fields. Adapting pretrained large language models offers transferable sequence representations, but tokenising spatial fields into 1D sequences can obscure spatial topology. Recent vision-language model (VLM) approaches encode images of hand-crafted statistics alongside fixed textual templates, leaving the correspondence between VLM architecture and ST structure largely unexploited. In this work, we propose VLM4ST, a framework that aligns the dual-path architecture of VLMs with spatial and temporal modelling. A 2D visual encoder processes spatial fields, while a 1D sequential encoder captures temporal dynamics. Multi-layer cross-attention couples both encoders, preserving spatial topology while leveraging pretrained sequence representations. To bridge the gap between natural image–text pretraining and ST data, we introduce Prompt Forge, a hierarchical pattern-memory module that learns latent ST prototypes and generates task-adaptive continuous prompts for cross-modal fusion. These prompts encode temporal periodicity, spatial correlations, and cross-dimensional dependencies. Extensive experiments show that VLM4ST outperforms task-specific, LLM-based, and recent VLM-based predictors, reducing short-term MAE by up to 16.5% (12.3% on average) relative to the strongest baseline on each dataset, with substantial gains in few-shot and zero-shot transfer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.