acceptodds
Under review as a conference paper at ICLR 2027

GeoSeg: Training-Free Reasoning-Driven Segmentation in Remote Sensing Imagery

Abstract

Recent advances in multimodal large language models (MLLMs) have reframed segmentation from fixed-category prediction toward open-ended, instruction-grounded localization. While reasoning-based segmentation has progressed rapidly in natural scenes, extending these advances to remote sensing remains an open challenge, owing to the scarcity and high annotation cost of reasoning-oriented training data, as well as domain-specific obstacles such as overhead viewpoints and extreme scale variation. We present GeoSeg, a training-free framework for reasoning-driven remote sensing segmentation that circumvents the supervision bottleneck by composing pretrained vision–language models for zero-shot inference. GeoSeg bridges MLLM reasoning and pixel-level localization through two key components: (i) bias-aware coordinate refinement, which corrects the systematic spatial offset introduced when MLLMs process overhead imagery, and (ii) dual-route prompting, which fuses visual keypoint cues with semantic text prompts for robust mask prediction. We also introduce GeoSeg-Bench, a diagnostic benchmark comprising 810 annotated image–query pairs across four scene domains and three hierarchical difficulty levels, designed to support standardized, zero-shot evaluation of reasoning-driven segmentation methods in remote sensing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.