StreamCast: Benchmarking MLLMs for Future Prediction in Interactive Livestreams
Abstract
Livestreaming is a real-time and interactive form of video, where events evolve continuously through visual, audio, and audience signals. Understanding livestreams requires MLLMs to move beyond retrospective comprehension and anticipate near-future outcomes from evolving context, a capability underexplored by existing benchmarks. In this paper, we introduce **StreamCast**, a benchmark for evaluating near-future prediction in livestreams. Each video contains an anchor question targeting at the prediction of the immediate future, with support questions probing capabilities relevant to the prediction: *perception localization*, *interaction dynamics*, *historical understanding*, and *causal attribution*. Constructed through hierarchical multimodal representation, agentic task formation, and automatic verification, StreamCast comprises 470 real livestreams across six categories and 1,421 questions. We further assess question quality and evidence support through human evaluation of a sampled subset. We evaluate 15 open-source and proprietary MLLMs and analyze performance across diagnostic dimensions, modalities, and languages. The strongest model reaches 63.19% accuracy on future prediction in the Full-context setting while achieving 85.07-90.67% across the four diagnostic dimensions, highlighting the difficulty of anticipating future developments even when prerequisite capabilities are relatively strong. We release our code and dataset at https://anonymous.4open.science/r/StreamCast.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.