TENSOR CLOCK: PUTTING ELAPSED TIME INSIDE VIDEO LANGUAGE MODELS
Abstract
Vision language models are now widely applied to video understanding. Questions about how fast an object moves or when an event happens, however, expose a gap: when frames are sampled non-uniformly, as a fixed frame budget requires, frame order no longer determines elapsed time, so identical frames can imply different speeds and the same motion can occur at different moments. Timestamps written into the prompt do not close this gap; a model trained with true frame-pair timestamps in its prompt still cannot tell apart rates that differ only in timing. This paper introduces Tensor Clock, an interface that lets elapsed time control the internal computation of a video language model rather than only appear in its input. Built on Qwen3.6-35B-A3B, it derives from the source timestamps of each frame pair the within-pair interval, the observation time, and the time elapsed since the previous block. These quantities set seconds-based rotary positions, drive elapsed-time decay in a measured subset of Gated DeltaNet recurrent heads, and condition the value features. Counterfactual training data, in which pixels are fixed and only the timing changes, force the answer to depend on time. Against a control trained with identical data, text timestamps, and schedule, Tensor Clock raises counterfactual rate accuracy from 0% to 95.8% and improves event-localization F1 by 0.07 and 0.03 on two event-timing tasks. Intervening on a trained model shows how the two carriers divide the work: rate comparison depends entirely on Tensor Clock, whereas event timing is carried by the text timestamps and further sharpened by Tensor Clock.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.