acceptodds
Under review as a conference paper at ICLR 2027

AwSAM: One-Shot Vocabulary-Free Segmentation of Anything with SAM-3

Abstract

Promptable vision foundation models (VFMs) such as SAM-3 can segment targets from text or geometric prompts, but each input creates distinct limitations for domain-specific applications. Text prompts fail for concepts outside the model’s learned semantic space, while geometric prompts incorporating points and boxes require specifying the target in every query image. Few-shot semantic segmentation (FSS) addresses this challenge by specifying a novel category at inference from a small set of annotated support examples. Existing FSS methods largely follow one of two routes: (i) converting support examples into geometric prompts by matching support and query patch features in a frozen VFM, which improves generalization but underperforms on fine-grained targets due to coarse patch features; or (ii) learning an explicit concept representation through task-specific training, tens of labeled images, or learned text-token conditioning, which improves performance but limits generalization. Given these limitations, we revisit FSS and SAM-3 prompt conditioning approaches and ask whether a single annotated example can be used to define a transferable concept directly, without training, a target dataset, or a class name to improve generalization? In our study, we demonstrate that a frozen SAM-3 can hold such a concept in its prompt space, a shared token sequence into which its text, geometric, and visual-exemplar encoders write. Fitting concept tokens in this space bypasses the text encoder entirely, and thus the language prior that limits generalization. We introduce AwSAM, a vocabulary-free FSS framework that learns a set of concept tokens from a single annotated example. We further study the contribution of external localization cues by pairing the fitted tokens with point prompts derived from a simple closed-form discriminant on VFM features. AwSAM sets a new state-of-the-art in one-shot vocabulary-free FSS across 12 diverse benchmarks, outperforming the strongest prior method by 11.3 mIoU while providing 2.9× faster query inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.