SURD-Bench: Benchmarking Unified Multimodal Models for Urban Spatiotemporal Cognition
Abstract
Understanding and modeling urban dynamics from Earth observation data is critical for urban remote sensing. Achieving this goal requires models to perceive urban objects in their spatial context, understand how they evolve over time, and predict their future states. Recent advances in Unified Multimodal Models (UMMs) have provided new opportunities to advance urban remote sensing from task-specific recognition toward comprehensive spatiotemporal cognition. However, existing benchmarks are mainly designed for general visual understanding or specific remote sensing tasks, and are not suitable for evaluating the unified understanding and generation capabilities of UMMs in urban remote sensing. The capability boundaries of UMMs in urban spatiotemporal cognition remain insufficiently characterized. Inspired by human cognitive processes, we propose SURD-Bench around four cognitive stages: Sensing, Understanding, Reasoning, and Decision, for evaluating UMMs in urban spatiotemporal cognition. In constructing SURD-Bench, we integrate in-house and public datasets, develop eight comprehensive tasks through an AI-assisted curation workflow with human verification, and design a unified evaluation framework for diverse outputs. Using SURD-Bench, we evaluate 18 representative commercial and open-source multimodal models. The results show that even the best-performing commercial model, Gemini-3.1-image, achieves an average score of only 41.85%, indicating that current UMMs still face substantial challenges in urban spatiotemporal cognition. Further analysis identifies limitations in fine-grained pixel-level semantic understanding, spatiotemporal relationship modeling, and the integration of understanding and generation capabilities. These findings offer insights into the future development of remote sensing UMMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.