acceptodds
Under review as a conference paper at ICLR 2027

Just Imagine Real-World Anomalies! Class-Specialized Vision-Free Training for Video Anomaly Detection and Understanding

Abstract

Real-world anomalies are difficult to collect: they are rare, unsafe, and often occur in privacy-sensitive environments, yet most video anomaly detection (VAD) methods require training videos. We ask: can a detector recognize anomalies it has only imagined through language? Building on evidence that text-trained detectors can transfer to video, we introduce NOVAD, a video-free framework for class-specialized anomaly detection and understanding. An LLM imagines normal and anomalous events as narratives ordered into beginning, following, climax, and resolution, with controlled variation over locations, objects, and participant counts. Encoded with a frozen vision–language text encoder, these narratives train a gated mixture of category-specialized experts, and their known temporal structure provides pseudo-labels for localization. At inference, the paired image encoder projects video frames into the same space. The predicted scores then guide multimodal event segmentation, and a video-language model explains only the resulting segments. Experiments on UCF-Crime, XD-Violence, and MSAD show that language-imagined events enable real-world anomaly localization without domain videos during detector training. On Vad-R1, detection-guided segments improve zero-shot reasoning over the untrimmed video, and a model trained with supervised fine-tuning only surpasses the full Vad-R1 model when fed with these segments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.