Rivet: Request-Independent Vision KV Cache for Zero-Recomputation VLM Serving
Abstract
Vision-language model (VLM) serving engines redundantly compute and store an image's decoder key-value (KV) cache for every request, even when multiple requests share the same image. Existing solutions fail to fully eliminate redundancy: prefix cache methods reuse KV only when text and positions align exactly, while position-independent approaches still require per-request recomputation, preventing a single physical cache from serving multiple requests. We present Rivet, which builds a request-independent vision KV cache that a serving engine computes once and shares across requests with zero recomputation. For VLMs with multi-dimensional rotary position embeddings (MRoPE), Rivet removes the dependence of vision KV on the request with (1) canonical MRoPE pinning, which replaces the position-dependent offset with a fixed canonical offset, producing identical post-RoPE keys across requests, and (2) sink-preserving isolation that retains the attention sink while blocking context-specific preceding text, making vision-token values request-independent. Together, these techniques produce a vision-KV object that is identical for every request, so one physical copy serves an entire batch through page-table sharing. When a long text precedes the image, Rivet shifts the text positions instead of the image and preserves the accuracy of full recomputation. On Qwen3-VL at 2B, 4B, and 8B scales, Rivet achieves accuracy on par with or even surpassing partial-recomputation baselines such as VLCache and CacheBlend across VQAv2, ChartQA, and TextVQA, while delivering up to 7.4 time-to-first-token speedup over full recomputation and 1.4–2.3 over previous methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.