Prediction Can Benefit Understanding: Visual Prediction as Complementary Supervision for Video LLMs
Abstract
Learning to predict how the visual world evolves is a powerful source of self-supervision, yielding strong video representations and world models that support planning and control. Video large language models (video LLMs), however, are trained almost entirely with language supervision; in the standard recipe, the video serves only as input and is never itself predicted. Prior work has unified visual generation and understanding in single models, yet whether learning to predict the visual world benefits understanding has not been systematically reported. In this work, we systematically investigate how visual prediction and understanding interact within a video LLM, asking whether learning to predict the visual world improves understanding. During model training, we remove a segment from a video and train the model to predict the visual features of that segment from the rest of the video and a description of the segment. A counterpart trained on the same data with language supervision alone serves as the control. At all three scales, from 0.8B to 4B, learning to predict improves every video understanding and temporal grounding benchmark we evaluate, with gains at 0.8B and 2B of up to 3.3 points on MVBench, 3.3 points on TempCompass and 3.4 points in mean grounding mIoU. Controlled ablations at multiple levels confirm that the gain is robust across design choices and data sources. These results show that visual prediction can serve as a general supervision paradigm for video LLMs, complementing language with a signal drawn from the visual world itself.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.