Anomaly-Guided Observation and Two-Level Semantic Feedback for Training-Free Video Anomaly Analysis
Abstract
Training-free language-based video anomaly detection (VAD) uses frozen vision– language models (VLMs) and large language models (LLMs) to localize abnormal events without task-specific model-weight optimization. Existing approaches in- creasingly exploit broader temporal and semantic context, yet generic context does not necessarily foreground abnormal events or specify how video-level seman- tics should modify local anomaly scores. We introduce a training-free semantic feedback framework that uses VAD-guided active observation for video anomaly understanding (VAU) and returns the resulting anomaly-focused description to refine temporal detection. Starting from an upstream anomaly-score curve, we adapt reasoning-based iterative frame selection: high-score regions seed the ob- servations, and anomaly scores together with VLM relevance assessments guide subsequent temporal exploration. A frozen video-language model generates a sin- gle open-ended anomaly description from the selected frames. The description is returned through two complementary feedback levels: video-level feedback de- termines whether the score distribution should be suppressed, preserved, or sharp- ened, with selective visual rechecking, while local feedback retrieves event-related windows and promotes scores only after target-local visual verification. All model weights remain frozen, and feedback is applied once without regenerating the de- scription. On the HIVAU video-summary evaluations of UCF-Crime and XD- Violence, VAD-guided observation improves GPT-rated reasonability, detail, and consistency over the matched CLIP-guided variant. For VAD, semantic feedback improves the re-evaluated upstream detector from 84.36% to 86.25% AUC on UCF-Crime and from 68.07% to 72.58% AP on XD-Violence, where the final AUC is 91.74%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.