acceptodds
Under review as a conference paper at ICLR 2027

Learning Lifetime-Aware Memory for Streaming Sports Commentary

Abstract

Streaming sports commentary requires models to identify participants, track unfolding actions, and decide when to describe an event using only the video observed so far. These decisions depend on interrelated semantic states with different useful lifetimes. Identity evidence remains relevant across events, active event states evolve as actions unfold, and completed events provide historical context. The challenge is to maintain these states over their respective lifetimes while preserving the semantic dependencies needed for accurate and timely commentary. In this paper, we propose GroundedPlay, a lifetime-aware memory framework that preserves identity evidence across events, updates active event states, and retains completed outputs as episodic context. It learns reciprocal interactions between identity and action representations, using identity evidence to refine action predictions and action context to select relevant player observations. In addition, context from the completion predictor guides evidence accumulation, while updated event states help predict whether the current event has ended. Experiments on NBA_Streaming and MatchTime-SID demonstrate that GroundedPlay outperforms state-of-the-art baselines in event response timing and commentary quality. It also achieves the highest video processing throughput among the evaluated methods. These results suggest that lifetime-aware memory and reciprocal state interactions offer a promising design approach for streaming tasks that require persistent entity knowledge, evolving event understanding, and timely responses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.