A Paradigm for Designing and Certifying Evaluation for Open-Ended Video Understanding
Abstract
Evaluating open-ended video understanding tasks is difficult because videos span diverse scenes and events, while valid responses vary in the information they convey and how they express it. Automated scoring can introduce errors that remain overlooked and obscure the capabilities being measured. We propose a paradigm for designing and certifying evaluation for these tasks. It combines a scoring procedure shared across evaluators with expert references and human baselines on a calibration panel. Certification requires an automated evaluation to match or exceed human-baseline agreement with the expert reference at every intermediate and final alignment point. Procedure design aims to reduce avoidable evaluation errors, while these comparisons make remaining disagreements visible. We instantiate the paradigm in detailed video captioning, assessing factual correctness and reference coverage through factor judgments and precision, recall, and F1 scores. In this instantiation, Gemini 3.7 Flash with agentic execution is the only certified individual evaluation among those tested. The paradigm also helps us compare and improve automated evaluations: averaging agentic outputs with those from a single scoring call to the same VLM improves alignment throughout the scoring procedure. Other evaluations exceed the human baseline in final-score correlation yet fail intermediate checks. These findings demonstrate how procedure-level analysis identifies evaluation weaknesses and guides further design and refinement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.