AsymmetricVAD: An Asymmetric Tri-modality Reasoning Paradigm for Weakly Supervised Video Anomaly Detection
Abstract
Weakly supervised video anomaly detection (WSVAD) is a critical task in intelligent surveillance. It aims to detect and localize abnormal events using only video-level labels. However, traditional multimodal WSVAD methods predominantly rely on symmetric fusion paradigms that treat video, audio, and text modalities with equal importance. Such paradigms often overlook representation biases, making the joint features highly susceptible to volatile background audio noise and temporal alignment offsets. To address these challenges, we analyze the role of different modalities and propose AsymmetricVAD, an Asymmetric Tri-modality Reasoning paradigm with a conditional loop. Specifically, AsymmetricVAD utilizes a Selective Injection and Enhancement (SIE) module as an asymmetric gatekeeper. It selectively injects audio into the video/text feature space only under temporal fluctuation and high cross-modal similarity. Furthermore, a Look-Back mechanism establishes the loop to suppress the representation bias. To justify our method, we build a new dataset TMVAD, which is the first tri-modality dataset. Extensive experiments on the benchmark XD-Violence and TMVAD datasets demonstrate that AsymmetricVAD achieves competitive performance on coarse- and fine-grained WSVAD tasks. We will open-source our code and dataset at https://anonymous.4open.science/r/AsymmetricVAD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.