acceptodds
Under review as a conference paper at ICLR 2027

Learning Video Camouflaged Object Detection from Unlabeled Static Images

Abstract

Video Camouflaged Object Detection (VCOD) is a criuital security technology for identifying camouflaged objects in videos. Existing VCOD methods often rely on motion signals across consecutive frames to capture dynamic cues. However, these motion signals are inherently noisy and inaccurate, and excessive reliance on them can lead to error accumulation. Moreover, the lack of camouflaged video datasets and the costly annotation process hinder progress in this field. To address these challenges, we propose SDC, a static-dynamic correspondence framework that extracts pseudo-dynamic signals from unlabeled static images, effectively capturing static-dynamic correspondences. Initially, we leverage coordinate data from the contrastive crop views to establish static correspondences, ensuring the consistency of static feature representations. Following that, we design a dynamic capture layer that generates pseudo-dynamic signals in both forward and backward directions. These pseudo-dynamic signals serve as cues for learning dynamic representations. Finally, we introduce a static-dynamic consistency loss function to enforce consistency between static and dynamic correspondences. Notably, our method requires only unlabeled static image data for training, freeing VCOD from its dependency on video data. By using only static images for training, SDC outperforms existing video-dependent strategies, surpassing state-of-the-art weakly supervised, unsupervised, and even surpasses several fully supervised video methods on public benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.