Video-HOCA: A Diagnostic Benchmark for Physical Anomaly Reasoning in Video-LLMs
Abstract
Video-HOCA evaluates how video-language models judge and explain visible physical anomalies. An operational Ontological-Causal taxonomy separates entity-centered from relation-centered violations, and four tasks on the same anomaly collection assess plausibility judgment, category attribution, fine-grained recognition, and mechanism explanation. The benchmark contains more than 1,439 videos and 3,470 question-answer pairs with human-verified labels and reference answers, plus a control of 250 generator-matched normal/anomalous pairs. Across 20 Instruct-mode models, nine-way category attribution remains difficult (mean Ontological/Causal macro-F1 of 37.9/46.0), and no model reaches the human reference on recognition or explanation. On the matched-generator control, the mean plausibility accuracy of 17 open-weight models is 64.4%, compared with 74.0% on the original real-versus-generated split; every model stays above chance, but the most accurate model on the original split is the least accurate on the control. For the tested model, shuffling frames barely changes plausibility accuracy yet lowers recognition and explanation scores. Annotation agreement and judge-sensitivity analyses characterize the reliability of these observations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.