Taming SAM3 for Few-shot Segmentation
Abstract
Few-shot semantic segmentation (FSS) requires segmenting target objects using a few reference images and masks. Conventional methods learn cross-image correspondences through episodic training, incurring training costs and posing challenges for generalizing to novel categories. A recent vision foundation model (VFM), SAM3, invites a different view: once a visual prompt identifies a concept, the model can propagate it to corresponding instances. This behavior, also known as promptable concept segmentation (PCS), mirrors FSS yet exposes a spatial mismatch: the support visual evidence and the geometry prompts inhabit different image coordinates. To this end, we introduce Visual Concept Transportation for SAM3 (VCT-SAM3), which brings the support concept into the query coordinate system to exploit SAM3’s native concept propagation. Rather than transferring the entire support scene, VCT-SAM3 extracts the annotated support object and places it into the query image. The inserted object then serves as a visual prompt, enabling SAM3 to discover and segment corresponding instances. Moreover, we explore a position-debiased semantic prior to protect potential target regions through the visual concept transportation. Our VCT-SAM3 uses a frozen SAM3, without auxiliary models or additional training. Experiments indicate that VCT-SAM3 achieves state-of-the-art (SOTA) performance (e.g., 7.9% mIoU gain on the COCO-20i benchmark) with lower latency and memory consumption. Our codes will be released to foster future research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.