Prune Once, Route per Turn: Constructing Persistent Visual Memory for Multi-Turn VLM Inference
Abstract
Large vision-language models (VLMs) have achieved impressive performance across diverse multimodal tasks, but their substantial inference cost becomes particularly challenging in multi-turn interaction, where the same visual input is repeatedly queried. Although visual token pruning provides an effective way to reduce inference cost, our empirical analysis reveals that naively extending single-turn pruning strategies to multi-turn scenarios yields two key limitations: cross-turn pruning mismatch, in which visual evidence discarded for one query becomes essential for subsequent turns; and costly query adaptation, where static retained subsets fail to accommodate evolving queries, whereas recomputing query-aware representations every turn introduces substantial latency overhead. Motivated by these findings, we introduce PersistRoute, a simple training-free framework that decouples session-level visual memory construction from turn-wise visual Key-Value (KV) routing. Specifically, PersistRoute comprises two complementary components: (1) Session Memory Prune (SMPrune), which conducts one-shot, query-agnostic visual-token pruning to build a diverse visual memory reused throughout the whole session, mitigating cross-turn pruning mismatch; and (2) Dynamic Query-Aware Visual KV Routing (DQVR), which selects query-relevant visual KVs from this persistent memory for each incoming query, enabling adaptive query processing without repeated pruning or memory reconstruction. Evaluated across multiple multi-turn benchmarks and diverse VLM backbones, PersistRoute establishes a more favorable performance-efficiency Pareto frontier compared to prior baselines. Using 50% persistent visual memory with merely 5% routed visual KVs, it preserves 96.2% of the full-model performance on MT-VQA-v2 and 95.6% on MT-GQA, delivering up to 2.54× end-to-end session-level inference speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.