Revisiting Co-Salient Object Detection via Explicit Semantic Consensus
Abstract
Co-salient object detection (CoSOD) aims to locate and segment the common salient object across a group of images. Existing methods infer group-level consensus in visual feature space and carry this implicit representation into object localization and segmentation. Yet visual similarity can cause semantically different distractors to enter the implicit consensus, reducing its reliability for segmentation. This motivates us to reconsider whether the essential shared information must remain encoded in a rich implicit visual representation, or can instead be preserved in a cleaner and more explicit form. We propose CoSEC (Co-Salient Explicit Consensus), a fully training-free framework that represents group-level consensus as a compact textual concept and uses it as the sole interface between group-level reasoning and per-image localization. A Semantic Consensus Estimator identifies the common object from the image group using a multi-image vision-language model, while a Consensus-Guided Segmenter independently grounds this concept in each image using a frozen text-conditioned segmenter. Without task-specific training or fine-tuning, CoSEC outperforms existing methods across all three benchmarks. On the challenging CoCA benchmark, it achieves an MAE of and an of , while requiring only s/image on a single GPU. These results suggest that once the shared semantic identity is reliably established, a compact explicit semantic consensus can replace a rich implicit visual representation for strong CoSOD performance. Code will be released soon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.