Prototype-driven Hierarchical Rule Induction and Matching for Efficient Video Anomaly Detection
Abstract
Video anomaly detection (VAD) aims to identify rare and unexpected events in video. Though training-free VAD using vision-language models (VLMs) has shown promise, existing pipelines rely on dense-frame semantic generation and reasoning, which is computationally expensive and not good at modeling context-dependent anomalies. To address these issues, we propose ProtoRule-VAD, an efficient framework that converts normal videos into reusable hierarchical semantic rules without anomalous supervision or model parameter updates. In the offline stage, scene-level and object-level normal prototypes are constructed from normal videos and used to induce contrastive hierarchical rules with a pretrained VLM, including scene-level, object-level, and condition-level rules. During inference, motion saliency and prototype-based visual novelty are combined to prioritize candidate segments. Condition-level rules then select relevant scene-level and object-level rules for selective rule matching and anomaly scoring. ProtoRule-VAD delivers competitive performance and interpretable rule-based evidence without requiring anomalous data and fine-tuning. Notably, our processing frame rate is over 500 FPS on one NVIDIA A6000 GPU, substantially improving online inference throughput over existing VLM-based VAD methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.