Cross-Domain Few-Shot Segmentation via Dual-level Distillation on SAM3
Abstract
Cross-Domain Few-Shot Segmentation (CD-FSS) aims to segment target objects in unseen domains using only a few annotated references. Recent Segment Anything Model (SAM)-based methods require less task-specific training but rely on pairwise visual matching, whose reliability decreases when reference and target appearances differ, leading to substantial variation in per-target IoU. SAM3 supports promptable concept segmentation and provides a semantic prompt space for learning a reusable category prototype. This prototype produces stable predictions across targets but cannot fully represent the different appearances. Based on these insights, we propose Dual-level Distillation on SAM3 (D-SAM3), a lightweight framework that uses self-distillation at complementary category and image levels. (i) Category Prototype Distillation transfers GT-box-prompted predictions into a reusable category prototype; (ii) Visual Residual Distillation calibrates reference-target visual correspondence separately for each category and expresses the resulting visual prediction as a residual relative to the prototype-based prediction; and (iii) Reliability-aware Fusion controls this residual based on correspondence reliability and uncertainty of the prototype-based prediction. Together, they produce a category-stable yet target-adaptive representation. With SAM3 frozen, D-SAM3 optimizes only approximately 8.2K parameters per category. Extensive experiments demonstrate strong CD-FSS performance, robustness to reference selection, and generalization across tasks and semantic granularities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.