acceptodds
Under review as a conference paper at ICLR 2027

Encoded but Unread: Locating and Recovering Within-Event Motion in Video Language Models

Abstract

Video large language models (Video-LLMs) answer questions about a video from sparsely sampled frames that a vision encoder turns into visual tokens for a language model to read. Questions about motion, such as how long an expression is held or whether a movement pauses, can only be answered by comparing what the model sees at different moments. Whether current Video-LLMs make this comparison, where the information is lost when they do not, and what recovery costs remain unclear. We introduce WarpPairs, a diagnostic benchmark of time-warp minimal pairs: each item re-indexes the frames of a real face, generated or action clip, so that positive and negative differ only in temporal structure and the label follows exactly from the edit. Six Video-LLMs perform at or near chance on all five edit types. The loss has two regimes. Below a threshold set by the temporal receptive field of a visual token, the sampler's frame stride times the encoder's temporal merge, a local scramble is attenuated for a linear probe on the tokens and invisible to the model, and the threshold moves with the architecture and with the sampling stride as predicted. Above the threshold the motion is recoverable from the tokens but absent from the answer: a linear classifier on difference statistics of the visual tokens reaches 72–90% balanced accuracy out of fold on three video sources, whereas the language model remains at chance. Supplying the missing comparison as temporal-difference (TD) tokens, computed from the model's own visual tokens, recovers the readout at a cost that depends on where the tokens enter and on the base model: at the decoder input they cut the fine-tuning steps to a fixed criterion up to six-fold at the same final accuracy, whereas the same signal added residually to the visual tokens is left at low weight by the model and saves nothing. The recovered readout transfers to unseen actors but only partly to other edits and sources. On controlled edits, these results separate a temporal-resolution scale set by sampling and encoding from a readout failure in the language model, and show that supplying the comparison saves most where frames-only fine-tuning stays longest in a constant-answer phase.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.