TennisRAM: Advantage-Aware Tennis Commentary through Joint Perception and State Modeling
Abstract
Advances in multimodal large language models (MLLMs) have extended tennis video understanding from event recognition to match analysis and commentary generation. However, high-speed ball motion, transient racket–ball contact, and continuous player interactions pose challenges for fine-grained perception and temporal understanding. Existing methods typically rely on specialized models to detect ball locations, player positions, and stroke events, and convert these predictions into structured representations for commentary generation. Although such representations provide explicit event cues, they often overlook continuous visual evidence, such as stroke preparation, body balance, and recovery dynamics. To address these limitations, we propose TennisRAM. We first adopt high-frame-rate video sampling to preserve rapid motion information. Then, predictions from auxiliary perception heads for ball localization, player pose estimation, and stroke recognition are used to guide spatial feature compression, improving computational efficiency while enhancing fine-grained player–ball interaction representations. To model within-rally dynamics, we further introduce a relative advantage state estimation task, where the model determines which player holds the advantage and provides supporting evidence using only the video observed up to the current time step. The estimated advantage states and their rationales are incorporated with visual features as conditions for commentary generation. For training and evaluation, we annotate 1,200 tennis rallies with relative advantage states and point-formation labels, yielding 4,322 advantage-state annotations. TennisRAM achieves 66.33% accuracy on relative advantage state estimation, with ROUGE-L and CIDEr scores of 33.11 and 44.15, respectively, for commentary generation. Ablation studies demonstrate that fine-grained perceptual guidance and conditioning on both advantage evolution and point outcomes improve commentary quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.