Auditing the Ground Truth: A Field-Aware Audit of Video Anomaly Understanding Benchmarks
Abstract
Language annotations in video anomaly understanding benchmarks guide model training and evaluation. Their quality must be assessed in relation to field responsibilities, temporal scopes, and granularity. We introduce a contract-aware audit framework that translates native field requirements into diagnostic questions and specifies the permitted evidence. Four tasks assess video quality, standalone text quality, native cross-text consistency, and video–text alignment. To determine whether a video–text pair meets multiple quality requirements simultaneously, we link video, standalone text, and alignment judgments to compute conditional joint usability. Our study includes 7,928 video entries and 56,103 video–text pairs across nine benchmarks. In the model-assisted audit, conditional joint usability ranges from 22.07% to 74.20% across benchmarks under field-macro aggregation. The predominant failure pattern is that video and text pass standalone checks, but their pairing fails alignment. To test whether audit-guided annotation revision improves performance, we revise CUVA descriptions and conduct controlled experiments with LVLMs, evaluating both description generation and multiple-choice question answering on VALU. The results show downstream benefits from revision, with effects varying by model and evaluation metric. Field-level audits further reveal quality differences within individual benchmarks. These diagnostics identify annotations requiring review and, together with native task coverage, guide the selection of benchmarks and fields suited to the evaluation goal and temporal granularity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.