ChronosKV: Dual-Clock Temporal Encoding for Streaming Video LLMs
Abstract
Streaming video LLMs process unbounded inputs by retaining a bounded visual cache and retrieving a sparse subset at question time. We demonstrate that this answer-time path suffers from a twofold information bottleneck: first, internal query-key attention locates the evidence little better than uniform sampling; second, standard position re-indexing preserves local token order while entirely collapsing absolute elapsed time. To resolve both limitations in a training-free framework, we introduce ChronosKV. ChronosKV pairs a lightweight contrastive vision-language retriever, delivering high-coverage evidence at constant query latency, with a dual-clock temporal encoding. By decoupling Rotary Position Embedding (RoPE) frequency bands band-by-band, ChronosKV routes high-frequency components to a compact ordinal track for local order, while steering low-frequency components to an absolute-time track for elapsed intervals. To evaluate what existing benchmarks miss, we introduce Interval-QA, a diagnostic suite isolating event order, recency, and relative metric distance. Across online streaming and long-video suites, ChronosKV attains the highest average accuracy of the compared systems, and restores the long-range elapsed time that previous position assignments discard.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.