acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Visual Evidence Extraction for Open-Vocabulary Video Visual Relationship Detection

Abstract

Open-vocabulary video visual relationship detection (Open-VidVRD) aims to recognize objects and their relationships in videos beyond a predefined vocabulary of object and relationship categories. Relationship classification typically relies on representations of trajectory pairs constructed from sampled frames, yet existing pipelines commonly divide videos or candidate relationship intervals into fixed-length temporal segments and sample the same number of frames from each segment. This uniform sampling strategy is suboptimal because relationships exhibit distinct temporal dynamics and require different numbers of frames to capture sufficient visual evidence. Stable relationships may require only a few frames, whereas dynamic ones require more. To address this issue, we propose an adaptive visual evidence extraction architecture for Open-VidVRD. It first partitions videos into temporal segments according to changes in optical flow and bounding box layouts, and then predicts the required number of sampled frames for each segment using aggregated appearance and motion cues. Subsequently, visual evidence is extracted from the selected frames and fused with contextual information from the entire segment for relationship classification. Due to the absence of annotations regarding the optimal number of sampled frames for each segment of a trajectory pair, we construct pseudo supervision by combining relationship-level temporal priors obtained from a multimodal large language model (MLLM) with annotated relationship durations and the temporal coverage of relationship instances within each segment. The MLLM is used only to derive offline supervision and is not invoked during inference. Extensive experiments on the VidVRD and VidOR datasets demonstrate that our method achieves state-of-the-art performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.