ReasonRSOS: Reasoning Salient Object Segmentation in Remote Sensing Images
Abstract
Remote sensing salient object segmentation (RSI-SOS) aims to segment the most visually salient objects in remote sensing images, and has achieved remarkable progress on conventional benchmarks. However, existing methods remain largely visual-driven and therefore lack the ability to interpret implicit user intent in complex geospatial scenes. To extend the research boundary, we introduce a new task, termed Reasoning Salient Object Segmentation in Remote Sensing Images (R-RSISOS). This task requires models to segment salient remote sensing objects from implicit text queries by leveraging geospatial context, complex reasoning, and world knowledge. To support this task, we construct ReasonRSOS, a benchmark containing 8,265 image-query-mask triplets. We further introduce IRSS, a challenging subset containing multi-salient-target scenes for high-level remote sensing reasoning.We also propose RISA, a Remote Instructed Salient Assistant that segments salient remote sensing objects according to implicit text queries. Extensive experiments show that RISA achieves superior performance on ReasonRSOS and demonstrates promising generalization ability on related tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.