acceptodds
Under review as a conference paper at ICLR 2027

ReMoM: Relation-aware Motion Matching for Motion-Guided Few-Shot Video Object Segmentation

Abstract

Motion-guided few-shot video object segmentation (M-FSVOS) aims to segment objects in a query video that exhibit motion patterns defined by a few support videos. A key challenge of M-FSVOS lies in support-defined motion correspondence, i.e., reliably matching motion across support and query videos despite variations in appearance and execution style. Although state-of-the-art video object segmentation models such as SAM 3 excel at object discovery, segmentation, and tracking, their object representations rely heavily on appearance cues and thus fall short in the M-FSVOS task. To tackle this challenge, we introduce a novel model named Relation-aware Motion Matching (ReMoM), which formulates M-FSVOS as a motion-centric masklet matching problem. Under this formulation, ReMoM performs support-conditioned motion matching through explicit track-to-way relation modeling over high-quality masklets extracted by SAM 3. Additionally, it strengthens the temporal evidence of the learned motion relations through counterfactual temporal enhancement and incorporates interaction-relevant context from surrounding objects to resolve interaction-dependent ambiguities. Extensive experiments on the MOVE and A2D-Motion benchmarks demonstrate the effectiveness of ReMoM, which significantly outperforms state-of-the-art models by 17.5% on MOVE.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.