Multi-modal Open-vocabulary Video Visual Relationship Detection
Abstract
Open-vocabulary Video Visual Relationship Detection (OV-VidVRD) aims to detect and localize subject-predicate-object relationship triplets in videos, even for relationships unseen during training. While pivotal for comprehensive video understanding, existing methods are largely confined to the visual modality. This uni-modal focus neglects rich contextual cues from other modalities, which are often vital for disambiguating complex relationships in real-world videos. To bridge this gap, we introduce the new task of Multi-modal Open-vocabulary Video Visual Relationship Detection (MM-OV-VidVRD) and present the first-of-its-kind benchmark VaM-VidVRD to catalyze research in this area. Our benchmark features numerous videos where each instance pairs the visual stream with a heterogeneous set of auxiliary modalities, such as audio, textual descriptions, and 3D data. Importantly, the benchmark encompasses many predicate categories that are visually indistinguishable yet semantically distinct—for instance, sing_to and speak_to share nearly identical visual appearances but are readily separable given audio cues—making auxiliary modalities essential rather than merely supplementary. This design mirrors the diverse and source-dependent nature of real-world multi-modal data. To tackle this new challenge, we propose a novel and flexible framework that effectively processes arbitrary combinations of modalities. The core of this framework is a Multi-modal Synergistic Prompting (MMSP) mechanism, which learns to dynamically align and fuse features from diverse sources into a unified and relationship-aware representation. Extensive experiments on our benchmark show that our method significantly outperforms vision-only approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.