DisCo: Distillation via Conflict for Weakly Supervised Video Anomaly Detection
Abstract
Video anomaly detection (VAD) has recently progressed along two promising directions: weakly supervised learning and adaptation of multimodal vision–language models (VLMs). Weakly supervised VAD is cost-effective because it relies only on video-level labels, but the lack of segment-level annotations makes it difficult to localize and disentangle multiple or diverse abnormal events. In contrast, VLMs pretrained on large-scale multimodal data encode rich, general semantic knowledge, yet they are not explicitly optimized for the domain-specific temporal reasoning required for anomaly detection. In this work, we propose a knowledge-sharing framework that transfers generalized semantics from VLMs to a weakly supervised VAD model via knowledge distillation. Empirically, we observe complementary behavior: zero-shot VLM predictions excel on certain anomaly categories (e.g., explosions and road accidents), whereas weakly supervised models are stronger on human-centric abnormalities (e.g., abuse). To enable selective distillation, we introduce a novel conflict score that quantifies class-wise disagreement by jointly modeling semantic and motion cues, and we distill knowledge only for classes with high conflict. Extensive experiments on two benchmark datasets show that our distillation via conflict (\bf DisCo) consistently improves anomaly detection performance over both weakly supervised and tuning-free VLM baselines, while requiring only a lightweight model at inference time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.