acceptodds
Under review as a conference paper at ICLR 2027

Boosting MLLM-SAM2 Referring Video Object Segmentation via Temporal Memory Reading

Abstract

Referring Video Object Segmentation (RVOS) aims to segment target objects in a video according to natural language expressions. Recent methods adopt an MLLM-SAM2 paradigm, where the MLLM token is projected into a sparse prompt for a SAM2 mask decoder. However, they suffer from object identification ambiguity when distinguishing visually similar instances via temporal cues. Through probing analysis, we trace this failure to the prompt interface: frame-order decodability degrades significantly during token-to-prompt conversion, yielding collapsed prompt representations that confuse same-category instances. We propose Temporal Memory Reading (TMR), which grounds expression-conditioned temporal queries in multi-frame SAM2 features and injects spatially adaptive query biases into memory attention throughout propagation. Experiments on multiple benchmarks show that TMR improves MLLM-SAM2 baselines on temporally ambiguous cases with similar distractors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.