acceptodds
Under review as a conference paper at ICLR 2027

BTGNet: A Background Text Guided Model with Joint Query Refinement for Few-Shot Semantic Segmentation

Abstract

Existing few-shot semantic segmentation (FSS) methods have extensively explored visual background information to suppress distracting regions, while multimodal approaches mainly rely on foreground-oriented text prompts. In contrast, the potential of background words as an explicit textual prior remains largely unexplored. We observe that category-related background words can provide complementary language cues for distinguishing target objects from visually confusing regions, while being easier to obtain and extend than image-specific background annotations. Based on this insight, we propose BTGNet, a Background Text Guided Network that explicitly incorporates background words into few-shot segmentation. Specifically, we employ a large language model to automatically generate task-adaptive, category-related background vocabularies and align them with visual representations through a frozen CLIP encoder to guide Grad-CAM-based query activation. To further improve the coarse activations produced by frozen CLIP, we propose a joint refinement strategy that combines fixed attention priors with learnable dynamic refinement. Extensive experiments on PASCAL-5\(^i\) and COCO-20\(^i\) demonstrate consistent improvements over state-of-the-art methods. Further analyses show that semantically relevant background words, rather than arbitrary additional text, are critical to performance gains, highlighting background words as a generalizable textual prior for multimodal FSS.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.