Map-Grounded World Model for Urban Street Generation
Abstract
World models have made remarkable progress in generating realistic and interactive environments, yet long-horizon urban generation remains difficult because both scene structure and travel direction must remain consistent as the viewpoint moves. Maps naturally provide these two forms of naturally: a map describes the spatial structure of an environment, while a route provides vectors that define how the viewpoint should move through it. We introduce Vision–Map–Language (VML), a map-grounded framework for urban street generation. Given a map and a route, VML uses the map's structure to determine what should appear around the viewpoint and route vectors to determine where the viewpoint should move next. These signals are translated into coordinated camera motion, language, and visual conditions for a pretrained video generator. VML further projects map structure into the current view to guide generation, allowing the map to remain active throughout long-horizon exploration. We evaluate VML through 3 capabilities: exploring a selected map location, following a specified route, and generating surrounding scenes that reflect the mapped structure along that route. Our results demonstrate that maps can serve as a simple and reusable representation for organizing and navigating generative urban worlds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.