COVSS: A Benchmark for Continual Open-Vocabulary Semantic Segmentation
Abstract
Semantic segmentation studies commonly evaluate continual updating and open-vocabulary recognition separately. Emerging work at their intersection leaves less understood how these capabilities develop together when supervision arrives incrementally. We introduce Continual Open-Vocabulary Semantic Segmentation (COVSS), a task and benchmark for studying this process through multiple class increments of comparable size. Built from existing COCO-Stuff171 and ADE150 annotations, COVSS specifies class partitions, stage-visible supervision, shared comparison conditions, and trajectory-based evaluation on internal and external vocabularies. A reproducible reference implementation combines anonymous scene supervision with dynamic semantic alignment. Experiments reveal weaknesses in both acquisition and retention among conventional continual segmentation methods, while open-vocabulary methods exhibit markedly different learning trajectories. Without dedicated anti-forgetting mechanisms, the COVSS baseline attains strong overall performance and positive net changes in previously learned categories, despite local declines. Pixel-level analysis attributes part of these gains to correcting earlier false positives through changes in competing category responses. These observations establish a benchmark for investigating how continual learning and open-vocabulary prediction interact, and identify challenges for developing scene understanding from successive, partially labeled experiences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.