OmniEarthSeg: Unified Instruction-Driven Segmentation for Heterogeneous Remote Sensing Imagery
Abstract
Multimodal large language models (MLLMs) enable instruction-driven segmentation of remote sensing (RS) imagery, allowing users to specify targets through natural language. However, existing approaches often rely on modality-specific preprocessing or fixed input configurations, limiting their ability to accommodate variable-channel observations. Their mask decoders also depend primarily on high-level semantic tokens, which provide limited explicit spatial cues for target localization. To address these challenges, we curate OmniGlobal-1.9M, a large-scale dataset comprising 1.9 million instruction-segmentation samples from over 800K globally distributed RS images acquired by 17 sensors and platforms, spanning RGB, multispectral, and synthetic aperture radar modalities. It provides diverse annotations, including multi-granular segmentation masks spanning a broad range of categories and long-form textual supervision that supports interpretable reasoning and response generation. Building on this dataset, we propose OmniEarthSeg, an MLLM that integrates the OmniSensor Tokenizer for unifying variable-channel inputs, the Latent Modality-Aware Modulator for enhancing the model's awareness of modality-specific characteristics, and Prompt-Guided Visual Enhancement for localizing task-relevant visual evidence before segmentation decoding. Extensive experiments demonstrate that OmniEarthSeg achieves state-of-the-art performance on six benchmarks spanning different sensing modalities and generalizes effectively to unseen datasets under zero-shot evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.