Diagnosing Target Preservation for Reliable MLLM-Based Referring Video Object Segmentation
Abstract
Referring video object segmentation requires identifying a language-specified target and preserving its identity throughout a video. However, mean segmentation accuracy alone provides limited insight into system reliability and the sources of failure. In this work, we investigate target preservation in a controlled multimodal large language model-based pipeline by tracing frame-level proposals, reference masks, and propagated predictions. We find that reference formation is a key bottleneck in this pipeline, even when the target is visible in the selected frame, and that propagation failures can persist despite ground-truth initialization. Guided by these findings, we propose a modular framework that separates video-level target resolution, single-frame reference formation, and video propagation. Support-aware proposal control withholds proposals when visual evidence is insufficient, while conditional recovery revisits these queries with additional observations. Experiments on MeViS demonstrate competitive segmentation accuracy, while controlled comparisons reveal trade-offs among output coverage, mask quality, and target-absent recall, highlighting the need to evaluate reliability beyond mean accuracy. The source code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.