acceptodds
Under review as a conference paper at ICLR 2027

CER-RVS: Contrastive Evidence Reflection for Agentic Referring Video Object Segmentation

Abstract

Referring Video Object Segmentation (RVOS) requires identifying and segmenting a language-referred object throughout a video, often based on complex cues involving actions, temporal relations, and object interactions. Recent training-free agentic approaches leverage the zero-shot reasoning capabilities of Multimodal Large Language Models (MLLMs) to avoid costly task-specific adaptation, yet remain limited by unreliable reasoning and imperfect object tracks. We propose CER-RVS, an evidence-verified, training-free agentic framework that jointly improves reasoning reliability and object-track quality. First, Contrastive Evidence Reflection replaces one-sided candidate pruning with explicit evidence competition: two reasoning agents independently collect supporting and counter-evidence for each candidate track, while an independent arbitrator contrasts their judgments, resolves conflicts, and provides targeted feedback for iterative revision. Second, Anchor-Guided Track Refinement selects reliable anchor masks based on mask reliability and temporal consistency, uses them to guide bidirectional refinement of candidate tracks, and incorporates identity-consistent refinements to improve mask quality and temporal completeness. Together, these components form a complementary perception–reasoning loop that strengthens visual evidence while enabling more verifiable and revisable reasoning. Extensive experiments on three challenging language-guided video object segmentation benchmarks demonstrate that CER-RVS establishes new state-of-the-art performance without additional training, substantially outperforming existing training-free approaches and surpassing representative supervised fine-tuning methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.