Look Inside Each Video: Rethinking How Video Anomaly Detection Is Evaluated
Abstract
Video anomaly detection (VAD) aims to localize anomalous events by identifying when they occur. A common evaluation pools frame-level anomaly scores across all test videos and measures whether anomalous frames receive higher scores than normal frames (Micro-AUROC). This pooling creates comparisons both within the same video and across different videos. However, the pooled evaluation protocol commonly used in current VAD benchmarks makes Micro-AUROC heavily dependent on cross-video comparisons, which do not directly test whether anomalous frames rank above normal frames within the same video. To quantify this imbalance, we decompose Micro-AUROC into within-video and cross-video comparisons. Across four widely used weakly supervised VAD benchmarks, within-video pairs account for only 0.071-0.388% of all anomalous-normal frame pairs. This imbalance can make Micro-AUROC a poor indicator of within-video ranking. Video-level classifiers that assign a single score to each video achieve 81.40-97.18% Micro-AUROC while obtaining 50% Within-AUROC. For existing weakly supervised VAD detectors, replacing all frame scores in each video with their mean sets Within-AUROC to 50%. Across 120 runs, this replacement retains a median 98.8% of the original Micro-AUROC margin above chance. Because the decomposition depends on the evaluation protocol rather than the training supervision, the same distinction applies whenever VAD frame scores are pooled across videos. High Micro-AUROC alone therefore does not establish temporal localization. Our analysis reveals a limitation of pooled VAD evaluation and motivates reporting within-video ranking alongside Micro-AUROC.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.