Towards Language-Guided Unified Visual Prompt Learning for Camouflaged Object Segmentation
Abstract
Camouflaged object segmentation (COS) aims to segment objects that seamlessly blend into the background, making it a crucial task in visual perception. Most existing COS methods formulate this problem in a purely visual manner, and thus struggle to capture the semantic cues and contextual knowledge required for understanding camouflage. Recently, language-guided segmentation methods have offered a promising solution. To advance language-guided COS, we study Language-Guided Unified Visual Prompt Learning for COS, where camouflaged objects are segmented under text-prompt guidance with diverse visual-prompt conditions, including mask, point, scribble, and null-prompt settings. To support this setting, we re-annotate existing COS benchmarks by collecting 12,487 text descriptions and constructing corresponding visual prompts, including full-mask, point, and scribble annotations, together with a no-prompt protocol. Building on this, we propose a unified COS framework based on a multimodal large language model (MLLM), which incorporates different visual prompts under text-prompt guidance. Extensive experiments show that our framework achieves competitive performance across different visual prompt settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.