acceptodds
Under review as a conference paper at ICLR 2027

IMAGINE IN DISCRETE SPACE: DISCRETE VISUAL LATENT REASONING FOR VIDEO EVENT PREDICTION

Abstract

Video event prediction requires multimodal large language models (MLLMs) to infer events beyond an observed video segment. We investigate whether future-video supervision can support this task through discrete visual intermediate states. We propose **Imagine in Discrete Space (IDS)**, which quantizes future-video features into visual tokens and learns to predict these tokens autoregressively from observed videos. A three-stage pipeline learns future-token prediction, grounds quantized future tokens in event descriptions, and refines event prediction with reinforcement learning. At inference, only the observed segment is available as visual input, and the final prediction is conditioned on self-generated visual tokens. The formulation provides a next-token training interface for future visual features without a pixel-reconstruction objective. We also introduce **VEPBench**, a benchmark of approximately 3K samples with recorded and plausible alternative futures and typed distractors, alongside a source-disjoint 17K-video training set. On VEPBench, IDS before reinforcement learning achieves 67.00% recorded-future selection, compared with 54.41% for the text-reasoning variant using the same backbone and reported SFT budget. The full model reaches 72.78%. Results on FutureBench provide evidence of cross-benchmark transfer, with stronger performance on one-hop prediction than on multi-hop tasks. These findings support discrete future visual tokens as a useful intermediate representation in the evaluated setting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.