OpenVA-Bench: Evaluating Open-World Video Anomaly Understanding
Abstract
Open-world video anomaly understanding requires models to recognize anomalies from known categories and make evidence-based inferences about anomalies outside the predefined category set. When available labels fail to describe an event, forcing it into a known category can produce an incorrect yet plausible interpretation. We introduce OpenVA-Bench to examine this risk through prediction accuracy, confidence reliability, and explanation quality. The benchmark integrates 1,330 videos from three datasets and constructs explicit Known–Unknown category splits, with Normal included in the Known set and Unknown Anomaly available as a rejection option. Across 25 Video-LLMs, stronger known-class recognition does not consistently coincide with reliable unknown rejection. Failures concentrate on semantically related category pairs, incorrect predictions can receive high confidence, and coherent explanations can conflict with reference event descriptions. Neither a larger model nor more input frames consistently yields better performance in the evaluated configurations. These findings expose a gap between recognizing known anomalies and providing reliable judgments in open-world applications. OpenVA-Bench supports systematic evaluation of this gap through prediction, confidence, and reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.