Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness
Abstract
Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC – under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, but their learned directions are approximately orthogonal to readiness and decode it less accurately than dedicated readiness probes. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to pp at matched latency with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.