ScaleCity: Joint Multi-Scale and Multi-Task Urban Region Representation Learning
Abstract
Urban region representation learning must accommodate tasks defined over different granularities and regional boundaries. Changes in these spatial configurations alter feature aggregation and target statistics, reflecting the scale and zoning effects of the Modifiable Areal Unit Problem (MAUP). We present ScaleCity, a multimodal urban region representation learning framework that jointly trains one model per city on mixed task–configuration samples. Joint training replaces separately fitted predictors with one model that supports multiple tasks and configurations. Graph-enhanced contrastive pretraining aligns street-view imagery, points-of-interest text, and remote-sensing features on atomic H3 cells. A prompt-conditioned hypernetwork adapts multimodal fusion and prediction to each task–configuration pair, while regional aggregation retains explicit extent information. Across three Chinese megacities, eight tasks, and nine configurations spanning grids, road-network partitions, and administrative regions, ScaleCity outperforms nine baselines in all task- and configuration-averaged R2 comparisons. Further experiments demonstrate zero-shot prediction at unseen configurations within the same city, while few-shot experiments examine adaptation to unseen tasks. These findings support joint multi-scale, multi-task learning of urban region representations. Code is available at https://github.com/far-rele/ScaleCity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.