CityGRASP: Agent Driven Geo-Reasoning with Adaptive Spatial Programming for Open-Vocabulary Urban 3D Grounding
Abstract
Open-vocabulary city-scale 3D grounding localizes language-specified objects and regions in large urban scenes beyond predefined categories. However, the appropriate search scope and geometric computation often become clear only as contextual references and local geometry are resolved, while premature filtering can exclude the target. We formulate city-scale grounding as adaptive spatial programming, in which spatial search and geometric computation evolve jointly with execution evidence and uncertain decisions remain recoverable. We present CityGRASP, an agent-driven geo-reasoning framework that instantiates this formulation. Starting from query-guided open-vocabulary entity memory, progressive scene–program co-planning adapts the search scope and subproblem order as contextual references are resolved. Scene-conditioned geo-code synthesis composes low-level point-cloud and geographic primitives into computations tailored to local geometry, returning execution evidence to guide subsequent planning. To correct premature exclusions, reversible grounding certification checks query requirements and competing hypotheses, directs evidence revision, and tracks dependencies to restore affected alternatives while reusing valid preceding computations. Experiments on CityRefer, CityAnchor, and CitySTAR-3D demonstrate effective target disambiguation amid visually similar objects and complex spatial relations, with absolute [email protected] gains of up to 17.73% over the state of the art.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.