acceptodds
Under review as a conference paper at ICLR 2027

AsymmetricVAD: An Asymmetric Tri-modality Reasoning Paradigm for Weakly Supervised Video Anomaly Detection

Abstract

Weakly supervised video anomaly detection (WSVAD) is a critical task in intelligent surveillance. It aims to detect and localize abnormal events using only video-level labels. However, traditional multimodal WSVAD methods predominantly rely on symmetric fusion paradigms that treat video, audio, and text modalities with equal importance. Such paradigms often overlook representation biases, making the joint features highly susceptible to volatile background audio noise and temporal alignment offsets. To address these challenges, we analyze the role of different modalities and propose AsymmetricVAD, an Asymmetric Tri-modality Reasoning paradigm with a conditional loop. Specifically, AsymmetricVAD utilizes a Selective Injection and Enhancement (SIE) module as an asymmetric gatekeeper. It selectively injects audio into the video/text feature space only under temporal fluctuation and high cross-modal similarity. Furthermore, a Look-Back mechanism establishes the loop to suppress the representation bias. To justify our method, we build a new dataset TMVAD, which is the first tri-modality dataset. Extensive experiments on the benchmark XD-Violence and TMVAD datasets demonstrate that AsymmetricVAD achieves competitive performance on coarse- and fine-grained WSVAD tasks. We will open-source our code and dataset at https://anonymous.4open.science/r/AsymmetricVAD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.