PACT: Pair-Aware Contextual Tokens for Video Visual Relation Detection
Abstract
Video Visual Relation Detection (VidVRD) focuses on detecting ⟨subject, predicate, object⟩ relations together with their temporal extents in videos. Most existing methods construct relation representations from features extracted from detected object regions. Consequently, contextual information from the surrounding scene is largely overlooked, and the representation remains dependent on the quality of the object detector. However, many predicates are determined by how two objects interact within the broader scene. We propose Pair-Aware Contextual Tokens (PACT), a relation representation that preserves global scene context while being aware of the target pair. PACT represents each frame using all patch tokens from a frozen image encoder. Since the patch tokens are shared across different pairs, PACT injects information about the target pair into them. Specifically, each token encodes its spatial overlap with the subject and object regions, together with the semantic embeddings of the two object categories. Experiments on ImageNet-VidVRD and VidOR show that PACT achieves state-of-the-art mAP performance on both benchmarks. These results show that preserving the global scene context while making the representation aware of the target pair is an effective way to learn relation representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.