acceptodds
Under review as a conference paper at ICLR 2027

MR2C: Multimodal Referring and Reasoning for Camouflaged Object Segmentation

Abstract

Camouflaged Object Segmentation (COS) has achieved significant progress in recent years, yet existing methods mainly focus on vision-only mask prediction and remain limited in understanding flexible user intent in interactive scenarios. To extend the boundaries of COS and facilitate future research in this field, we propose Multimodal Referring and Reasoning Camouflaged Object Segmentation (MRR-COS), a new task that aims to segment the queried camouflaged object according to multimodal user references. To support this task, we construct MR2C, a multimodal benchmark containing 9932 images, 64 categories, 9932 segmentation masks, 19864 textual references, 1920 reference images, and 640 reference sounds. MR2C stands out with three key features: (1) four types of multimodal references, including referring text, reasoning text, reference images, and reference sounds; (2) an emphasis on understanding user intent and target-related sound content rather than merely detecting camouflaged regions; and (3) the incorporation of complex reasoning and world knowledge into its textual references. Furthermore, we introduce MCOSA, a multimodal camouflaged object segmentation assistant, to address multimodal understanding, reasoning, and fine-grained segmentation in MR2C. Extensive experiments show that MCOSA outperforms existing methods on MR2C, demonstrating the necessity of the proposed task and the effectiveness of our benchmark and framework. Our code and dataset will be publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.