Seeing Twice: Hierarchical Re-Observation for Training-Free Video Anomaly Detection
Abstract
Training-free video anomaly detection relies on large pretrained models to detect abnormal events without task-specific training. However, most existing methods inspect model-generated descriptions only once, making them vulnerable to implicit and ambiguous anomaly cues. Specifically, context-dependent anomaly cues may remain hidden at a single temporal granularity and become discernible only through multi-granularity observation, while semantically ambiguous cues may persist when the initial visual description is incomplete or unreliable. In this paper, we propose Seeing Twice, a hierarchical re-observation framework that uncovers implicit cues at the observation level and resolves semantic ambiguity at the reasoning level. Our method constructs a hierarchical event representation with intervals at multiple temporal granularities. At the observation level, to recover implicit cues, the framework selectively revisits intervals insufficiently observed at a single granularity, recovering context-dependent evidence by jointly examining them with their hierarchical companions. At the reasoning level, to resolve semantic ambiguity, the framework selects between the initial and recovered descriptions according to their agreement with the visual evidence. It then reconciles the corresponding anomaly scores through hierarchical score propagation to produce coherent frame-level predictions. Seeing Twice achieves 87.12% ROC-AUC on UCF-Crime and 75.93% AP on XD-Violence, surpassing existing methods by a clear margin. These results demonstrate the effectiveness of the proposed framework for training-free video anomaly detection. Moreover, our method provides interpretable, human-readable cues for anomalous event analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.