UniRefDet: Universal Referring Detection of Salient and Camouflaged Objects in Unconstrained Scenes
Abstract
Salient object detection (SOD) and camouflaged object detection (COD) with language-driven understanding have attracted substantial attention in recent years. Existing studies predominantly focus on referential foreground object localization under individual task settings, while language-guided distinction between salient and camouflaged objects, particularly when these two visual attributes coexist in the same scene, remains underexplored. In this paper, we construct UniRefDet, a novel referring segmentation benchmark that unifies salient and camouflaged object detection tasks in unconstrained real-world scenes. The UniRefDet presents fine-grained referring expressions that explicitly describe object-environment relations. Built on this benchmark, we propose COREFNET, a SAM-based multimodal segmentation framework. We first propose a Language-Aware Semantic Adaptation (LASA), which inserts lightweight Text Adapters into the text encoder to adapt concept-level representations to fine-grained referring expressions. An Attribute Discrimination Module (ADM) is further designed, aiming to calibrate multiple proposals within each attribute and suppresses conflicting salient and camouflaged responses in coexistence scenes. Extensive experiments show that COREFNET achieves state-of-the-art performance across all scene configurations on UNIREFDET.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.