RoadBench 2.0: Benchmarking Multimodal Agents on End-to-End Urban Road Network Construction for Traffic Simulation
Abstract
Large-scale remote-sensing imagery and tool-using agents make automated urban road-network construction increasingly feasible. However, a simulation-ready network must maintain consistency across junction localization, corridor connectivity, road direction, lane structure, and lane-to-lane movements; existing benchmarks typically evaluate these tasks in isolation and do not reveal how errors accumulate through a construction chain. We introduce RoadBench 2.0, a dataset and evaluation framework for urban road-network construction from remote-sensing imagery, with aligned OSM references. The dataset contains 166 regional cases covering six US cities, and organizes structural construction and SUMO simulation checks into a connected evaluation protocol, including both stage-wise evaluation with controlled inputs and end-to-end evaluation that passes model predictions downstream. We compare five representative models under a unified protocol. The results show that models can perform well on selected individual stages but remain unreliable across the complete chain, with the highest end-to-end success rate reaching only 0.066. Failure analysis shows that errors in junction localization, corridor connection, and simulation checks accumulate downstream, while ablations reveal stage-dependent effects of OSM. These findings identify cross-stage consistency, error recovery, and simulation feedback as central challenges for simulation-ready road-network construction, and establish RoadBench 2.0 as a unified basis for studying them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.