EvoCDS: Unifying Video Context-Dependent Concept Segmentation via Evolving Concept Discriminative Subspaces
Abstract
Unified image context-dependent (ICD) concept segmentation has made substantial progress, yet extending it to video requires more than temporal modeling. Across natural videos, medical videos, and 3D medical scenarios, video context-dependent (VCD) concept segmentation must handle heterogeneous concepts whose discriminative evidence evolves with the visual context, together with severe cross-task imbalance in video and frame counts. We characterize this temporal evolution as concept drift and propose EvoCDS, the first unified framework specifically designed for VCD concept segmentation. Motivated by the compact low-dimensional structure of concept-specific discriminative information, EvoCDS constructs an evolving concept discriminative subspace that preserves reliable historical directions while adapting to current evidence. We further introduce temporal clip balancing to stabilize joint learning across heterogeneous sequence datasets. Across eight tasks spanning natural and medical scenarios, EvoCDS consistently outperforms existing unified models and achieves competitive performance against specialized methods. We also construct Video-CD, a benchmark of 201 videos and 34,562 frames for evaluating robustness to concept drift. Together, EvoCDS and Video-CD provide a unified framework and benchmark for studying evolving CD concepts in dynamic visual scenarios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.