Ref-AVIS: Referring Audio-Visual Instance Segmentation in Long Videos
Abstract
In real-world audio-visual scenes, people use language to specify particular objects among multiple candidates. Existing benchmarks, however, provide limited support for independently segmenting and tracking these objects in complex, long videos. We introduce Ref-AVIS, a benchmark for referring audio-visual instance segmentation comprising 1,187 videos across 26 sound-source categories, with an average duration of 59.88 seconds. It covers scenes with multiple objects of the same or different categories and provides diverse referring expressions, independent masks, and persistent identities. We further develop a baseline that incorporates video-level referring understanding into instance segmentation through competitive instance grounding, audio-guided adaptation, and audio-visual discrimination. These designs enable independent segmentation and consistent tracking of multiple referred objects in dynamic audio-visual scenes. Extensive quantitative and qualitative experiments show that accurate foreground segmentation does not necessarily translate into reliable instance tracking, while demonstrating our baseline’s improvements in both segmentation and tracking on Ref-AVIS. Together, Ref-AVIS and our baseline support the development and evaluation of multimodal models for language-guided instance segmentation and tracking in long videos. Our dataset and baseline code will be made publicly available upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.