acceptodds
Under review as a conference paper at ICLR 2027

MAVER: Multi-Anchor Verification and Evidence Recovery for Training-Free Video Reasoning Segmentation

Abstract

Reasoning video object segmentation identifies and segments objects from implicit queries about actions, relations, or temporal events. Although training-free methods use multimodal large language models (MLLMs) to interpret such queries, a plausible interpretation can still yield a coherent track of the wrong object. In this paper, we propose MAVER (Multi-Anchor Verification and Evidence Recovery), a training-free framework that treats inferred object descriptions as hypotheses and checks candidate regions against the original query before using them as tracking prompts. Multi-anchor verification searches candidate frames under a fixed hypothesis, providing alternative grounding opportunities. Early stopping ends the search at the first confirmed anchor and uses its frame and box as the single initial prompt. When an entire grounding round confirms no anchor, evidence recovery revises the hypothesis and repeats grounding within a preset round limit, keeping the original query unchanged. If verification remains inconclusive, memory-based fallback can supply an explicitly unconfirmed prompt from the retained candidates. MAVER coordinates frozen reasoning, detection, and segmentation models without task-specific fine-tuning. Extensive experiments on four referring and reasoning video segmentation benchmarks demonstrate that MAVER achieves the best overall performance among the compared methods. Moreover, MAVER improves performance while maintaining inference efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.