Just Imagine Real-World Anomalies! Class-Specialized Vision-Free Training for Video Anomaly Detection and Understanding
Abstract
Real-world anomalies are difficult to collect: they are rare, unsafe, and often occur in privacy-sensitive environments, yet most video anomaly detection (VAD) methods require training videos. We ask: can a detector recognize anomalies it has only imagined through language? Building on evidence that text-trained detectors can transfer to video, we introduce NOVAD, a video-free framework for class-specialized anomaly detection and understanding. An LLM imagines normal and anomalous events as narratives ordered into beginning, following, climax, and resolution, with controlled variation over locations, objects, and participant counts. Encoded with a frozen vision–language text encoder, these narratives train a gated mixture of category-specialized experts, and their known temporal structure provides pseudo-labels for localization. At inference, the paired image encoder projects video frames into the same space. The predicted scores then guide multimodal event segmentation, and a video-language model explains only the resulting segments. Experiments on UCF-Crime, XD-Violence, and MSAD show that language-imagined events enable real-world anomaly localization without domain videos during detector training. On Vad-R1, detection-guided segments improve zero-shot reasoning over the untrimmed video, and a model trained with supervised fine-tuning only surpasses the full Vad-R1 model when fed with these segments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.