Benchmarking Multimodal Self-Evolving Memory for Personalized Vision-Language Agents
Abstract
Personalized AI assistants need long-term memory that evolves as users interact with them over time and across domains. Yet evaluations focused on text and static memory snapshots reveal little about how multimodal memory supports lifelong personalization. We introduce MemoryVCD, a benchmark that separates three memory design dimensions: representation, updating, and growth. Built on real users’ time-stamped multimodal histories, it combines five input formats with two growth axes: within-domain progression with advancing prediction targets, and cross-domain expansion with fixed targets. We evaluate six personalization tasks using up to twelve vision-language backbones from three providers and four main memory-update methods. Visual information generally improves personalization, but no update method performs best across all tasks. Accumulated history helps ranking and choice, while compression and recency can better preserve quantitative accuracy and local detail. More history does not always improve performance, and cross-domain memory can cause persistent negative transfer in visual prediction. These findings point to memory systems that selectively retain and transfer evidence as users evolve. MemoryVCD provides a shared testbed for studying multimodal lifelong personalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.