Who Does What to Whom? Role Binding for Interaction-Aware Video Segmentation
Abstract
Interaction-aware referring video object segmentation (InterRVOS) asks not only which objects are referred to, but who does what to whom: the model must segment the Actor and Target of a described interaction according to their asymmetric semantic roles. Unlike conventional referring segmentation, these roles are not intrinsic object identities but are defined by the interaction between participants. Existing methods use separate Actor and Target tokens while leaving this interaction implicit, which can lead to ambiguous or relationally inconsistent role predictions. We instead formulate InterRVOS as relation-conditioned role binding, where the interaction serves as an explicit intermediate state that connects multimodal understanding to role-specific segmentation. Based on this view, we introduce R3Refer, which first learns an input-dependent, visually grounded interaction representation. It then performs directed role reasoning, where the interaction conditions Actor prediction and the resulting Actor representation further guides Target reasoning. Finally, triadic refinement jointly coordinates the Interaction, Actor, and Target representations before mask decoding. Experiments on InterRVOS-127K show consistent improvements in both Actor and Target segmentation, while evaluations on conventional RVOS benchmarks demonstrate the broader effectiveness of relation-aware role modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.