Discover and Verify: Fine-Grained Video Object Segmentation with Weak-Text Prompts
Abstract
Language-guided video segmentation converts text queries into spatial prompts, reducing the need for manual spatial annotations. However, this paradigm commonly relies on an implicit assumption that semantic similarity is sufficient to determine the identity of the target. This assumption often fails in fine-grained settings where only an easily available category name is provided as a weak-text prompt. Visually similar objects from different categories may receive similar semantic similarity scores, making it difficult to reliably distinguish the target from confusing objects. To address this issue, we propose a selective weak-text prompted fine-grained segmentation framework based on a “discover-and-verify” mechanism. Confuser-aware Prompt Discriminative Grounding (CPDG) compares the target category with potential confuser categories to learn latent discriminative cues. These cues are then used to locate and aggregate local discriminative evidence within candidate regions, allowing CPDG to rerank the candidate regions. Semantic Consistency Verification (SCV) further combines semantic and visual evidence to verify the top-ranked candidate and rejects the candidate hypothesis when the available evidence is insufficient. Extensive experiments on two weak-text fine-grained datasets, EndoVis17 and EndoVis18, demonstrate the effectiveness of our method, which consistently improves weak-text prompted fine-grained segmentation performance and improves mcIoU, our primary evaluation metric, over the strongest compared method by 1.31% and 16.91%, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.