IMAGINE IN DISCRETE SPACE: DISCRETE VISUAL LATENT REASONING FOR VIDEO EVENT PREDICTION
Abstract
Video event prediction requires multimodal large language models (MLLMs) to infer events beyond an observed video segment. We investigate whether future-video supervision can support this task through discrete visual intermediate states. We propose **Imagine in Discrete Space (IDS)**, which quantizes future-video features into visual tokens and learns to predict these tokens autoregressively from observed videos. A three-stage pipeline learns future-token prediction, grounds quantized future tokens in event descriptions, and refines event prediction with reinforcement learning. At inference, only the observed segment is available as visual input, and the final prediction is conditioned on self-generated visual tokens. The formulation provides a next-token training interface for future visual features without a pixel-reconstruction objective. We also introduce **VEPBench**, a benchmark of approximately 3K samples with recorded and plausible alternative futures and typed distractors, alongside a source-disjoint 17K-video training set. On VEPBench, IDS before reinforcement learning achieves 67.00% recorded-future selection, compared with 54.41% for the text-reasoning variant using the same backbone and reported SFT budget. The full model reaches 72.78%. Results on FutureBench provide evidence of cross-benchmark transfer, with stronger performance on one-hop prediction than on multi-hop tasks. These findings support discrete future visual tokens as a useful intermediate representation in the evaluated setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.