VPCache: Answer-Preserving Cache-View Selection for Repeated-Visual LVLM Inference
Abstract
Repeated-visual LVLM inference answers multiple questions about the same image, document, chart, or screenshot. Exact prefix reuse saves repeated encoding but retains the full visual representation, while compression can discard evidence needed by a later question. We propose VPCache, which selects reusable cache views according to paired answer sufficiency. Its Answer-Preserving Cache View Selector (AP-CVS) qualifies candidate views within question cells, prioritizes answer quality, and freezes the resulting routes before evaluation. A learned high- fidelity view, APVC-hifi, refines a deterministic anchor for reuse across questions. On 448 held-out requests from seven workloads with Qwen2.5-VL-7B-Instruct, VPCache adds two soft and two exact matches over the guarded base, with no paired soft losses against either the base or full-prefix inference. Static compression controls incur 25–72 paired soft losses against the base. Across all held-out requests, VPCache reduces injected visual tokens by 4.90% and selected-view memory slots by 2.46% on average, matching the base’s representation budget while improving answer quality. These results support question-cell selection as a way to govern reusable visual compression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.