Boosting MLLM-SAM2 Referring Video Object Segmentation via Temporal Memory Reading
Abstract
Referring Video Object Segmentation (RVOS) aims to segment target objects in a video according to natural language expressions. Recent methods adopt an MLLM-SAM2 paradigm, where the MLLM token is projected into a sparse prompt for a SAM2 mask decoder. However, they suffer from object identification ambiguity when distinguishing visually similar instances via temporal cues. Through probing analysis, we trace this failure to the prompt interface: frame-order decodability degrades significantly during token-to-prompt conversion, yielding collapsed prompt representations that confuse same-category instances. We propose Temporal Memory Reading (TMR), which grounds expression-conditioned temporal queries in multi-frame SAM2 features and injects spatially adaptive query biases into memory attention throughout propagation. Experiments on multiple benchmarks show that TMR improves MLLM-SAM2 baselines on temporally ambiguous cases with similar distractors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.