LoomKV: Structured Low-Rank KV Caches for Video LLMs
Abstract
Among the most accurate ways to compress a video language model’s KV cache, a low-rank factorization shared across a group of layers must first hold the very cache it compresses. We compute it per input during prefill, keeping a fixed-rank representation rather than a fixed fraction of tokens. Which structure to keep is itself a result: a matched-storage control rejected our first, grid-preserving Tucker design in favour of sharing across layers (47.97 against 59.12), and the grid does not pay for itself after that extraction either. The offline cross-layer factorization, however, consumes the completed prompt cache, which is exactly what cannot be held. LoomKV removes that dependence. As visual blocks arrive it maintains compact coefficients, a shared channel basis and the Gram matrix needed to refresh that basis, and a refresh from r + Sb rows is exact for the whole history the state represents, so the dense multi-layer cache is never materialized. On a 24 GB card the offline factorization fails at 128 frames and a chunked dense control at 160; LoomKV completes 224, per-session resident state is 5.7 to 7.7× smaller, and nine prefilled sessions stay resident against one. At a matched answer-time cache it outperforms InfiniPot-V by 7.46 points on Qwen2-VL and 13.56 on Qwen2.5-VL on MME-VideoOCR, stays within 3.6 points of the offline ceiling across three model families, and on long Video-MME is indistinguishable from dense while retaining 11.9% of the visual cache. Controlled ablations attribute the gain to cross-layer sharing; refresh frequency, coefficient transport and causal feedback have no additional cost this sample resolves. All primary comparisons are paired, and negative results are reported.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.