Cardano: Training Streaming Video Language Models to Know When They Have Seen Enough
Abstract
Video language models (VLMs) answer a question about a video after processing all of it, so inference cost scales with video length rather than with the difficulty of the question. Processing the whole video is also not always beneficial, since additional video can overturn an answer that was already correct. The relevant question is therefore not how much video a model watches but when it should stop. We present Cardano, a training formulation for streaming VLMs that learns to answer questions and to decide when further observation would not change the answer. Rather than predicting the content of unseen video, Cardano predicts whether observing more video would change its current answer, which makes adaptive stopping a self-supervised problem: the target is obtained by comparing the model's answer on a partial video with its own answer on the full video, and requires no annotation or dataset-specific supervision. At inference the model reads the video window by window and, after each window, either commits to an answer or continues watching. A commit head of about forty parameters, fit once and deployed at its natural operating point with no threshold tuned on any dataset, outperforms a causal fixed budget at a matched number of frames on four of six benchmarks, with confidence intervals excluding zero, and transfers without refitting to a second architecture. Stopping preserves accuracy rather than improving it: with Qwen2.5-VL-7B, Cardano is within 1.5 points of its whole-video accuracy or above it on all six benchmarks while processing between one third and four fifths of each video. On FunQA and Video-MME, where training the answerer on video prefixes raises accuracy, Cardano exceeds the untrained backbone watching the whole video by 8.4 and 3.8 points. Code is available at https://anonymous.4open.science/r/cardano-vlm-404A.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.