GLSC: Global-Local Semantic-Structural Collaboration for Semi-Supervised Remote Sensing Segmentation
Abstract
Vision-language models and self-supervised vision foundation models provide complementary priors for semi-supervised remote sensing segmentation: CLIP offers category-level semantic knowledge, whereas DINOv3 preserves fine-grained spatial structures. However, effectively coordinating these heterogeneous priors at both category and spatial levels remains challenging. We propose GLSC, a global–local semantic–structural collaboration framework that strengthens the interaction between CLIP and DINOv3 through two complementary modules. First, Shared-Query Response Fusion (SQRF) uses CLIP text prototypes and learnable DINOv3 queries to project the two visual streams into category-conditioned response spaces and fuses their responses while preserving an explicit category dimension. Second, Dual-Context Response Refinement (DCRR) refines the fused responses by modeling intra-category spatial dependencies under visual-feature guidance and inter-category relationships under query guidance. The refined responses are subsequently aggregated and reintegrated into the original visual representations through residual enhancement. During training, an asymmetric auxiliary guidance loss encourages DINOv3 to learn from CLIP, while gradients from this auxiliary pathway are stopped at the CLIP branch. Experiments on Potsdam, LoveDA, and WHDLD under four labeling ratios show that GLSC outperforms the reported Co2S results in all 12 dataset–ratio settings, achieving an average absolute improvement of 1.14 mIoU points and a maximum gain of 2.11 points on LoveDA with 1/8 labeled data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.