Observe, Hypothesize, and Verify: Active Evidence-Seeking Reasoning for Training-Free Video Anomaly Detection
Abstract
Video Anomaly Detection (VAD) aims to identify rare and complex anomalous events in videos, which requires the model to capture fine-grained details and reason about their contextual relationships. However, existing training-free VAD methods commonly follow a passive caption-and-judge paradigm, where video clips are first converted into generic textual descriptions and anomaly reasoning is restricted to evidence already captured by generic video descriptions via vision-language models. While effective for visually salient events, this paradigm is insufficient for evidence-demanding anomalies, which depend on subtle object interactions, temporally distributed clues, and careful verification. Therefore, we propose OHV-VAD, which formulates video anomaly detection as an observe-hypothesize-verify reasoning process for active evidence seeking and verification. This paradigm is built around three evidence-centered capabilities: 1) Hypothesis-Driven Evidence Acquisition transforms passive captioning into targeted perception by using diagnostic questions as evidence probes to acquire anomaly-relevant visual evidence. 2) Adaptive Temporal Evidence Refinement maintains a compact evidence state through a clue memory that continuously organizes anomaly-relevant clues across video segments. 3) Reflective Anomaly-Normality Verification evaluates each suspicious segment under competing abnormal and normal explanations, mitigating over-detection caused by one-sided anomaly reasoning. Extensive experiments on VAD benchmarks demonstrate strong overall performance among training-free approaches. Category-wise analyses further show that OHV-VAD effectively detects evidence-demanding anomalies while retaining strong performance on visually salient events, and false-positive analyses confirm its ability to reduce over-detection of normal behaviors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.