acceptodds
Under review as a conference paper at ICLR 2027

DISP: Decoupled Image-Semantic Prompting with Multi-scale Visual Alignment for Remote Sensing Referring Segmentation

Abstract

In remote sensing referring segmentation, models must localize socially defined regions—such as parks, schools, and commercial zones—whose boundaries depend not only on visual appearance but also on functional and contextual semantics. This makes socio-semantic segmentation especially challenging, since visually similar regions may correspond to different social meanings and require fine-grained cross-modal reasoning. Existing vision-language segmentation paradigms still face two key limitations: single-token methods entangle semantic reasoning and spatial localization within one representation, while multi-stage VLM-to-SAM pipelines compress rich socio-semantic understanding into coarse geometric prompts; moreover, they insufficiently model multi-scale regional visual statistics needed for precise masks. To address these challenges, we propose DISP (Decoupled Image-Semantic Prompting with Multi-scale Visual Alignment), a unified framework that decouples the language-semantic role and visual-spatial role into two dedicated tokens, [SEG] and [IMG]. These tokens are projected through separate pathways and used as independent sparse prompts for the SAM2 decoder under a split-path dual-image architecture, enabling collaborative yet disentangled semantic-spatial segmentation. We further introduce Multi-scale Masked Feature Matching (MS-MFM), a training-only supervision module that aligns the [IMG] token with attention-pooled target-region representations from multiple ViT layers, allowing it to capture hierarchical visual cues with zero inference overhead. Experiments on socio-semantic segmentation benchmarks demonstrate that DISP achieves consistent SOTA performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.