acceptodds
Under review as a conference paper at ICLR 2027

Multimodal Influence Learning for Urban Spatio-Temporal Forecasting: A Cross-City Foundation Model and Dataset

Abstract

Urban spatio-temporal foundation models learn reusable forecasting patterns by pretraining across cities, yet they rely mainly on single-modal spatiotemporal observational data. Multimodal information such as roads, remote sensing, weather, and events provides predictive evidence beyond these observations, but transferring the learned use of multimodal information across cities remains an open problem. A further unaddressed challenge is that modalities contain heterogeneous content, and individually beneficial modalities can interfere when used jointly. To this end, we propose MILU, a multimodal urban spatio-temporal foundation model that learns each modality's predictive influence rather than directly mixing modality content. MILU encodes each modality separately, retrieves historical spatio-temporal context conditioned on both the forecasting query and modality information, and composes the resulting influences in a shared influence space to modulate forecasting states under bounded scaling. MILU is jointly pretrained across cities by comparing forecasting errors with and without multimodal information, so that it learns how to use multimodal information under varying availability. This transfer also requires pretraining data with diverse modalities, ample observations, and broad geographic coverage, which existing urban foundation model resources do not yet provide. We therefore construct the CityModal Dataset, covering 87 regions across four continents with approximately 19.21 billion valid spatio-temporal data points, and associate urban observations with road, remote-sensing, weather, and event information. Experiments show that, without target training, MILU achieves the lowest MAE on 28 of 30 short-term and 27 of 30 long-term forecasting tasks, and achieves increasing gains as more modalities are used jointly. With one tenth of the target training data, it outperforms the strongest task-specific baseline trained on the full target set in aggregate performance. Controlled comparisons further show that multimodal utilization transfers to new scenarios, extends to modality combinations unseen during training, and improves with limited target-data adaptation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.