When, Where and What to Comment: A Clue-grounded Benchmark for Sports Video Commentary
Abstract
Real-world sports commentary requires more than generating descriptions from video clips. A competent commentator must continuously determine when an event deserves commentary, identify where the supporting visual evidence occurs, and decide what to say based on the event and its competitive significance. Existing benchmarks largely reduce this process to text generation from pre-segmented clips, implicitly specifying commentary timing and weakening the need for fine-grained visual grounding. We introduce CG-SC-Bench, a clue-grounded benchmark for sports commentary that evaluates these three capabilities in a unified framework. Each commentary instance is associated with explicit visual clues, enabling the benchmark to assess whether a model identifies commentary-worthy moments (When), localizes the evidence supporting the commentary (Where), and generates factually consistent commentary conditioned on the observed clues (What). CG-SC-Bench covers 8 major sports categories and 27 subcategories, comprising 2,684 commentary segments with fine-grained clue annotations. Based on these annotations, we construct three complementary evaluation tasks for commentary timing, visual clue localization, and clue-conditioned commentary generation. We evaluate a broad range of closed- and open-source vision-language models on CG-SC-Bench. The results reveal substantial limitations of current models in temporal decision-making, fine-grained visual grounding, and evidence-grounded factual generation, showing that fluent sports descriptions do not necessarily translate into reliable continuous commentary. CG-SC-Bench provides a diagnostic evaluation framework for developing sports video models that can comment at the right time, ground their statements in the right evidence and say the right thing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.