VastFSS: Vast Category Few-Shot Segmentation
Abstract
In this paper, we propose a novel benchmark, dubbed **VastFSS**, aiming to facilitate the development of more generalizable few-shot segmentation (FSS) by covering a vast number of categories across diverse scenarios. Specifically, VastFSS contains 104K images from 2,400 classes, spanning a wide range of application scenarios, from everyday natural scenes to specialized domains, such as biology, agriculture, medical imaging, remote sensing, aerospace science, and other scientific imagery. Compared to existing datasets, our VastFSS provides a broader and more diverse platform for training and evaluating generalizable FSS models, thereby facilitating their real-world applications. In addition, to encourage future research, we present the ***M****ultimodal large language model (MLLM)*-***A****ssisted* ***FSS***, termed **MA-FSS**, a novel framework which leverages the visual understanding capabilities of MLLMs for improving FSS. The key idea is to introduce a set of learnable object tokens to extract target-relevant representations from the support sample via an MLLM and employ them as additional prompts for a segmentation model to predict the query mask. To enhance the representation, we present mask-guided attention within the MLLM to direct the object tokens towards foreground target in the support image, together with a simple adaptive token gate to dynamically adjust their importance, leading to better performance. In extensive experiments on VastFSS, we show that MA-FSS achieves promising performance, outperforming other FSS methods, and generalizes well in a zero-shot manner to existing FSS benchmarks, demonstrating both the effectiveness of MA-FSS and benefits of the broad category and scenario coverage in VastFSS. Our benchmark and code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.