More Multimodality Is Better? Uncovering the Functional Roles of Multimodal Information in Agent Memory
Abstract
Multimodal large language model (MLLM) agents increasingly rely on multimodal memory to store and reuse past experiences for long-horizon tasks, yet it remains unclear when, where, and how multimodal information benefits memory. We present the first systematic study of how different levels of multimodal integration within the memory pipeline affect agent performance. We establish a four-level multimodal memory hierarchy that progressively introduces multimodal capabilities into memory construction, storage and retrieval, and utilization, ranging from text-only to fully multimodal memory. Experiments are conducted under long-horizon settings across 4 representative memory-intensive task domains, 10 benchmarks, and 4 MLLM backbones. Our results reveal that deeper multimodal integration does not consistently improve performance. Instead, its benefits depend on whether newly introduced multimodal capabilities address remaining evidence gaps. Three critical roles of multimodal information are identified: a Semantic Enricher during memory construction, an Experience Locator during retrieval, and an Evidence Resolver during utilization—capturing distinct ways multimodal information supports downstream decision-making. Deeper integration, however, may also introduce distraction and yield limited or negative returns. These findings challenge the prevailing assumption that richer multimodal memories universally benefit agent performance and highlight the importance of stage-aware multimodal memory design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.