acceptodds
Under review as a conference paper at ICLR 2027

Cross Granularity Decoupled Decoder and Temporal Query Evolution for Referring Video Object Segmentation

Abstract

Referring Video Object Segmentation (RVOS) aims to segment the target specified by a language expression throughout a video. Recent diffusion based RVOS methods benefit from text aligned spatiotemporal representations, but converting diffusion features into accurate and temporally stable referring masks remains challenging. Frame wise query decoding may suffer from target drift under occlusion, disappearance, and similar distractors. Meanwhile, a unified mask prediction head has to handle both target localization and complete mask generation in the same prediction space. To address these limitations, we propose a prediction stage adaptation framework for diffusion based RVOS. First, we introduce Temporal Query Evolution (TQE), which updates object queries with short term temporal context and text aligned long range references to maintain target identity over time. Second, we design a Cross Granularity Decoupled Decoder (CGD), which decomposes mask prediction into discriminative localization and semantic coverage. The Discriminative Instance Branch performs query conditioned instance filtering to identify the referred target, while the Holistic Semantic Branch provides broader semantic activation for mask completion. We employ Temporal Context Mask Refinement (TCMR) to refine low resolution masks with temporal boundary cues and improve high resolution mask quality. Experiments on five RVOS benchmarks validate the framework's effectiveness and generalization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.