SFC-SAM3: Specializing Foundation Priors via Cross-Modal Coordination for Referring Image Segmentation in Agricultural Remote Sensing
Abstract
Referring image segmentation (RIS) grounds free-form language in pixel-level masks, but extending it to agricultural remote sensing exposes a substantial mismatch between generic foundation-model priors and domain-specific visual and linguistic structure. Agricultural scenes contain large and weakly bounded regions, while referring expressions combine phenology, morphology, and spatial relations. To address these challenges, we introduce SFC-SAM3, a parameter-efficient foundation-prior framework for agricultural referring remote sensing image segmentation (RRSIS) that specializes complementary structural, visual, and linguistic priors. Rather than retraining foundation backbones, SFC-SAM3 coordinates SAM3, DINOv3, and CLIP through a frequency-aware PromptGenerator, a multi-layer DINOv3 FPN bridge, a CLIP token-level text bridge, and a residual mask refinement head, together with LoRA residuals on the SAM3 encoder. Training follows two stages: joint fine-tuning of the adapters and SAM3 mask decoder, followed by parameter-efficient continued training with the decoder frozen. Across the two-stage training procedure, SFC-SAM3 updates only 4.3M parameters (0.35% of the 1.21B model). On FarmSeg-VL, SFC-SAM3 reaches 87.04% mIoU, compared with 86.11% for RMSIN trained on the same data. The advantage becomes clearer off the training distribution: without training on the target datasets, SFC-SAM3 reaches 41.14% mIoU on the RRSIS-D land subset and 19.81% on the RefSegRS land/ground subset, compared with 5.65% and 13.99%, respectively, for the strongest evaluated baselines. These results point to a practical lesson for agricultural RRSIS: task-structured specialization of complementary foundation priors can preserve strong in-domain performance while substantially improving transfer to unseen land-oriented remote-sensing scenes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.