VCRM: Empowering Efficient Memory for Vision-Language Models with Visual Context Record Memory
Abstract
Although large language models excel at complex reasoning tasks, effectively leveraging continuously growing long-term historical information remains a critical challenge. Existing memory-augmented methods primarily rely on storing and retrieving historical information in textual form. However, as interaction history continues to grow, memory content faces a dual bottleneck of ever-increasing context overhead and key evidence dispersed across voluminous retrieval results. To address these issues, we propose Visual Context Record Memory (VCRM), which achieves significant token compression with stable performance gains. VCRM consists of two synergistic components: (1) Hierarchical Evidence Prospecting (HEP), which filters redundant information through tiered confidence routing and adaptive cliff pruning; (2) Discriminative Semantic Rendering (DSR), which renders critical evidence into two-dimensional visual records with non-uniform information density, enabling key evidence to be preserved more effectively under constrained token budgets. We evaluate our method on three widely adopted benchmarks spanning long-term conversational memory and multi-hop question answering, i.e., LoCoMo, LongMemEval-S, and HotpotQA. Results show that VCRM improves downstream task performance while substantially reducing token consumption, validating the effectiveness of the method itself. Under constrained budgets, VCRM maintains near-full performance through visual memory, demonstrating strong robustness under compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.