CityExpander: Agent-Guided Synchronized Diffusion for Large-Scale City Generation
Abstract
Generating large-scale urban semantic layouts from text requires preserving both city-wide spatial coherence and fine-scale structural detail while satisfying spatial planning constraints. Directly generating high-resolution semantic image over large urban areas is computationally expensive, while lowering the resolution can erase narrow roads and merge nearby buildings. Sliding-window generation limits per-window computation but introduces sequential dependencies and may produce seams or repetitive patterns. Meanwhile, directly conditioning a generative model on text provides only coarse control over the locations, proportions, and spatial relations of land-use categories. We present CityExpander, a text-guided diffusion framework that addresses both challenges. First, We introduce a compiler-grounded condition plan agent that translates natural-language descriptions into executable zone-ratio conditions. To balance computational cost and structural fidelity, we use a global DiT for coarse city structure and a local DiT for fine-scale refinement. We propose semantic consensus globally synchronized denoising (SC-GSD) to achieve semantic consistent layout generation with simultaneous denoising, avoiding the sequential dependency of sliding-window sampling. Extensive experiments demonstrate that CityExpander effectively generates high-quality text-guided city layouts at both single-tile and large spatial scales. Further evaluations validate the effectiveness of SC-GSD in promoting semantic consistency and reducing time cost during large-scale synthesis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.