NFL-World: A Multimodal World Model for Understanding American Football Plays
Abstract
We propose NFL-World, a multimodal world model for American football that learns compact, predictive game-state representations from synchronized video and player-tracking data. NFL-World compresses each modality into latent tokens and combines them through a causal Transformer that autoregressively predicts subsequent modality tokens, capturing both player interactions and evolving play dynamics. We first learn modality-specific representations and then train predictive fusion, allowing the model to build on informative states from both streams. Experiments show that NFL-World outperforms evaluated baselines in overall play classification, next-second player-motion forecasting, and similar-play retrieval. Further analyzes show that video complements tracking to improve play understanding and tactical retrieval, while compression substantially reduces the fusion token budget and preserves or improves play understanding performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.