Basketball-QA: Diagnosing Perception and Reasoning in Multi-Player Basketball Video Understanding
Abstract
Multimodal large language models (MLLMs) show increasingly strong capabilities in sports video analysis, including question answering, commentary generation, and event recognition. However, existing evaluations rarely combinemultiple interacting players, dense sequences of events, and fine-grained spatialand temporal relations within a single video understanding task. We introduceBasketball-QA, a basketball video understanding benchmark with over 8K questions covering player matching, spatial grounding, action recognition, temporalgrounding, event relations, counting, hallucination and tactical analysis. Solving these questions requires models to jointly identify relevant players, determine theiractions and locations, and associate events across time. In end-to-end evaluation of video-input MLLMs, the best multiple-choice accuracy is 76.0% and the best open-ended score is 2.87/5, revealing substantial limitations in complex multi-player basketball scenes. We then localize the source of these failures througha series of increasingly targeted experiments. Probing atomic perception testsshow that errors already arise when recovering the underlying facts. When correct perceptual facts are instead provided in text, a small language model without any task-specific fine-tuning achieves 78.5% multiple-choice accuracy, surpassing the best video-input model we evaluate. Together, these experiments indicatethat current MLLMs are primarily limited by the acquisition and association offine-grained visual evidence on Basketball-QA, rather than by reasoning over correctly provided perceptual information. We hope Basketball-QA will support thesystematic evaluation and improvement of MLLMs for complex sports video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.