3D-VL-Fly: A Multi-task 3D Vision–Language Benchmark for Aerial Post-Disaster Damage Assessment
Abstract
Recent advances in multimodal large language models have enabled sophisticated visual understanding and reasoning for disaster response. Emerging benchmarks such as DisasterM3 have begun to extend multimodal learning to post-disaster scenarios through tasks such as visual question answering and damage assessment. However, these benchmarks predominantly rely on 2D imagery, limiting the evaluation of multimodal models' ability to reason over the three-dimensional geometric structure of disaster environments. Important evidence, including collapsed roof structures, debris extent, building-level spatial relationships, and partially occluded structures, are difficult to fully characterize from a single 2D view. In this study, we introduce 3D-VL-Fly, the first 3D vision–language benchmark specifically designed for post-disaster damage assessment. Built from 3D point-cloud scenes generated from aerial imagery collected following Hurricane Ian (2022), the benchmark provides building-level instance segmentation in addition to semantic segmentation of roads and trees, attributes per building independently, and question–answer pairs spanning eight question families with spatially grounded building references. Beyond visual question answering, 3D-VL-Fly incorporates both referring segmentation, which evaluates a model's ability to localize language-referred buildings or regions, and reasoning segmentation, which evaluates whether a model can identify the 3D evidence relevant to answering a reasoning-oriented question. This unified formulation enables the evaluation of multimodal models not only on what they answer, but also on where the relevant evidence is located in a 3D disaster scene. We evaluate a range of pretrained 3D and 2D vision–language models and encoders, and fine-tune a representative 3D vision–language model. Released models remain below a simple majority baseline on VQA and achieve low segmentation accuracy, indicating substantial room for improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.