MetricKV: Decoded-Space Geometry for Visual KV Compression in Streaming Video Understanding
Abstract
Streaming video models observe visual evidence before future queries are known, making compact storage of historical visual key–value (KV) states important for long-running inference. We introduce METRICKV, a query-agnostic visual-KV codec derived from the decoded-KV reconstruction objective. For any positive-definite quadratic decoded-space distortion and full-column-rank decoder, decoded error decomposes exactly into a continuous projection term and an additional finite-rate latent error under a decoder-induced pullback metric. This metric determines the same ordering over finite-rate candidates as the decoded-KV objective; absorbing it into the latent coordinates yields a canonical Euclidean coding space in which standard nearest-neighbor quantization is exactly decoded-error aligned. Within this canonical family, minimizing the coder-independent continuous projection term reduces to PCA in the distortion-transformed KV space. We instantiate the framework with Layer-RMS reconstruction geometry and progressive residual vector quantization (RVQ). The experimental results demonstrate that a single progressive RVQ stream achieves dynamic-KV compression while retaining more than 98% of FP16 performance on both OVO-Bench and StreamingBench. Code will be released soon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.