Rethinking Interactive Image Segmentation: A Prototype Selection View
Abstract
Interactive image segmentation aims to extract target objects based on the user guidance, and recent click-based approaches gain popularity for their simplicity and efficiency. However, existing deep interactive methods often struggle to retain sparse spatial prompts, capture intra-class variance, and leverage label semantics beyond predefined channels. In this paper, we rethink interactive segmentation from a prototype learning perspective, where each user click defines a prototype comprising both spatial and label attributes. We propose a novel framework that learns spatial prototypes to guide feature flow and progressive label prototypes to model object characteristics, enhancing the embedding of user intention and enabling parallel segmentation querying without fixed label constraints. Instead of direct selecting an object, the joint representation of multiple prototypes helps to capture the intra-instance variance. Additionally, we introduce difficulty-aware and similarity-driven training mechanisms to enhance focus on challenging regions and improve inter-class discrimination. Vast experiments demonstrate that our method outperforms current state-of-the-art methods across diverse benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.