OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
Abstract
Although multimodal large language models have shown strong potential for remote sensing disaster assessment and reasoning, most existing methods still follow a direct prediction paradigm that maps multimodal inputs to final answers, lacking an explicit and verifiable intermediate reasoning process aligned with pixel-level visual evidence. This limitation reduces the interpretability and reliability of their outputs. To address this issue, we propose the Disaster Vision-Language Reasoner (DisasterVLR), a unified vision-language reasoning model for remote sensing disaster assessment that improves fine-grained change understanding and reliable reasoning through an explicit intermediate thinking process and multi-task balanced reinforcement learning. DisasterVLR is built on two key techniques. First, self-evolving chain-of-thought generation constructs visual-evidence-constrained seed reasoning trajectories through multi-agent collaboration and iteratively expands them to build a high-quality Disaster Assessment and Reasoning (DAR) dataset. Second, ARSPO++ (Adaptive Reward Scaling Policy Optimization Plus) dynamically adjusts optimization signals through intra-task reward mapping and inter-task weight regulation, alleviating training imbalances caused by differences in task difficulty. Experimental results demonstrate that DisasterVLR achieves superior performance across a comprehensive range of disaster-related tasks, exhibiting stronger overall disaster assessment capabilities than existing natural-scene multimodal models and remote sensing vision-language models. Compared with the strongest competing method, DisasterVLR reduces WAPE by – percentage points on damage assessment tasks, including damaged-building counting and damaged-road area estimation, and improves mIoU by points on tasks such as referring expression segmentation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.