acceptodds
Under review as a conference paper at ICLR 2027

GLSC: Global-Local Semantic-Structural Collaboration for Semi-Supervised Remote Sensing Segmentation

Abstract

Vision-language models and self-supervised vision foundation models provide complementary priors for semi-supervised remote sensing segmentation: CLIP offers category-level semantic knowledge, whereas DINOv3 preserves fine-grained spatial structures. However, effectively coordinating these heterogeneous priors at both category and spatial levels remains challenging. We propose GLSC, a global–local semantic–structural collaboration framework that strengthens the interaction between CLIP and DINOv3 through two complementary modules. First, Shared-Query Response Fusion (SQRF) uses CLIP text prototypes and learnable DINOv3 queries to project the two visual streams into category-conditioned response spaces and fuses their responses while preserving an explicit category dimension. Second, Dual-Context Response Refinement (DCRR) refines the fused responses by modeling intra-category spatial dependencies under visual-feature guidance and inter-category relationships under query guidance. The refined responses are subsequently aggregated and reintegrated into the original visual representations through residual enhancement. During training, an asymmetric auxiliary guidance loss encourages DINOv3 to learn from CLIP, while gradients from this auxiliary pathway are stopped at the CLIP branch. Experiments on Potsdam, LoveDA, and WHDLD under four labeling ratios show that GLSC outperforms the reported Co2S results in all 12 dataset–ratio settings, achieving an average absolute improvement of 1.14 mIoU points and a maximum gain of 2.11 points on LoveDA with 1/8 labeled data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.