Understanding Before Forecasting: Multi-City Urban Dynamics Prediction with Vision-Language Model
Abstract
Multi-city urban dynamics prediction requires capturing fine-grained spatial-temporal patterns across heterogeneous cities while effectively transferring knowledge across diverse urban environments. Vision-language models (VLMs), with their large-scale pretraining, multimodal representation, and generalization capabilities, offer a potential opportunity for this task. However, VLMs lack domain-specific knowledge of spatial-temporal urban dynamics and an understanding of flexible prediction tasks that vary in target dynamics, spatial regions and scales, and prediction times. Moreover, heterogeneous city distributions can induce conflicting optimization objectives during joint learning, hindering effective cross-city knowledge sharing. In this paper, we propose UrbanBind, a unified visual-language framework for multi-city urban dynamics prediction with three tightly coupled components: (i) QA-Guided Spatial-Temporal Knowledge Grounding, which adapts a pretrained VLM to urban dynamics and flexible prediction-task semantics through structured question-answering supervision; (ii) Multi-City Urban Dynamics Encoding, which explicitly learns spatial-temporal representations across cities, complementing QA-guided grounding with spatial-temporal knowledge and task understanding; and (iii) Nash-Balanced Multi-City Optimization, which balances conflicting city-wise updates to shared parameters while preserving city-specific optimization. Extensive experiments on multiple real-world urban dynamics datasets demonstrate the effectiveness of UrbanBind for multi-city forecasting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.