LifeKV: Full-Lifecycle KV Cache Retrieval for Streaming Video Question-Answering
Abstract
Streaming VideoQA suffers from increasing GPU memory consumption and inference latency as the visual context grows. Existing KV retrieval methods reduce this overhead, but typically target only specific inference stages and lack a unified mechanism for controlling memory and latency throughout streaming inference. We propose LifeKV, a training-free full-lifecycle KV retrieval framework that keeps historical visual KV states revisitable during visual prefilling, question prefilling, and autoregressive decoding. The complete historical visual KV memory is stored in RAM, while only a compact query-relevant context is maintained on the GPU. LifeKV further introduces layer-ahead asynchronous KV prefetching to overlap descriptor-based retrieval and RAM-to-GPU KV recall with ongoing Transformer computation. Experiments show that LifeKV maintains competitive VideoQA accuracy while substantially reducing the growth of GPU memory usage and inference latency as the visual context length increases, providing an efficient and scalable solution for streaming video understanding over existing VideoQA models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.