acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Target Preservation for Reliable MLLM-Based Referring Video Object Segmentation

Abstract

Referring video object segmentation requires identifying a language-specified target and preserving its identity throughout a video. However, mean segmentation accuracy alone provides limited insight into system reliability and the sources of failure. In this work, we investigate target preservation in a controlled multimodal large language model-based pipeline by tracing frame-level proposals, reference masks, and propagated predictions. We find that reference formation is a key bottleneck in this pipeline, even when the target is visible in the selected frame, and that propagation failures can persist despite ground-truth initialization. Guided by these findings, we propose a modular framework that separates video-level target resolution, single-frame reference formation, and video propagation. Support-aware proposal control withholds proposals when visual evidence is insufficient, while conditional recovery revisits these queries with additional observations. Experiments on MeViS demonstrate competitive segmentation accuracy, while controlled comparisons reveal trade-offs among output coverage, mask quality, and target-absent recall, highlighting the need to evaluate reliability beyond mean accuracy. The source code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.