Rethinking Open-World Video Anomaly Detection: Diagnosing Definition Blindness
Abstract
Open-world video anomaly detection (OWVAD) aims to localize anomalous events in videos according to a user-specified definition. An OWVAD detector should therefore use the supplied definition to decide what is anomalous. We call this behavior definition following, and its failure definition blindness. However, existing temporal evaluation protocols may obscure this distinction, making strong generic anomaly detection appear to reflect definition following. For example, Drift@5 and the customizable VAD protocol treat both normal frames and anomalies outside the current definition as negatives. Across the three benchmarks used in our study, normal frames account for 78.6%-96.4% of these negatives. As a result, the reported score can be driven largely by target-versus-normal separation rather than target-versus-other-anomaly separation. Consistent with this diagnosis, a LaGoVAD branch that never receives the anomaly definition nearly matches the full model under Drift@5. We therefore introduce three complementary definition-conditioned evaluations that separate definition following from generic anomaly detection. Applying them to VAD systems and general vision-language models on UCF-Crime, XD-Violence, and MSAD reveals that strong performance under existing evaluation can coexist with weak definition following. Motivated by these results, we introduce Definition-Contrastive Scoring (DeCoS), which reduces anomaly evidence shared across competing definitions and improves definition following across the three benchmarks. Our results show that generic anomaly detection and definition following are distinct capabilities and should be evaluated separately in OWVAD.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.