The Decoder Is the Temporal Model: Verdict-First Reasoning for Online Video Anomaly Understanding
Abstract
Accurately localizing anomalous moments in videos typically relies on additional classification heads or temporal detection modules, which are effective for fast and precise anomaly localization. However, in scenarios that also require semantic interpretability, existing approaches often introduce separate detection components, making it difficult to fully exploit the inherent video understanding and language reasoning capabilities of vision-language models (VLMs). We observe that the causal context modeling and fixed vocabulary projection of a language model can directly provide a continuous online anomaly score. Based on this insight, we propose Vigil, a VLM-native framework for online video anomaly detection that directly turns a vision-language model into a window-level anomaly detector while preserving its capability to analyze anomalous events. To accommodate streaming-video scenarios, Vigil optimizes the prefill stage of LLM inference and jointly encodes historical frames and current-window frames into a unified model context, without introducing an additional temporal aggregator or recurrent cross-window hidden state. During supervised fine-tuning and inference, the model predicts Normal or Anomaly at the first generation position, and the corresponding token logits are normalized to produce a continuous anomaly score, thereby eliminating the need for a separate detection head. Under a predefined 20-second context setting, Vigil achieves low-latency real-time inference while outperforming the strongest offline baselines on both UCF-Crime and XD-Violence. These results suggest a simple and reproducible pathway toward video anomaly detection systems that retain both efficient online detection and semantic understanding, particularly for edge deployment scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.