MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
Abstract
Long-term agent memory is becoming increasingly multimodal, but existing evaluations rarely test whether agents can preserve visual details across long interactions. In many prior benchmarks, questions involving images can still be answered from captions, dialogue, or other textual cues, without relying on the original visual information. Existing evaluations also provide limited coverage of cases where visual states change over time and later evidence updates or replaces earlier observations. We introduce MemEye, a benchmark built around two challenges in multimodal memory: preserving visual information at the right level of detail and using that information correctly as memory evolves over time. We capture these challenges with two axes, ranging from scene-level to pixel-level visual evidence, and from single-evidence retrieval to reasoning over changing states. Under this framework, we construct a new benchmark of 1,200 questions across eight life-scenario tasks, with 100 questions per taxonomy cell and ablation-driven validation gates for assessing answerability, shortcut resistance, visual necessity, and reasoning structure. We evaluate 13 memory methods across four vision-language model (VLM) backbones and find that current systems . Further analysis shows that failures often come from lost visual details, stale or incomplete retrieval, and difficulty tracking the current visual state.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.