acceptodds
Under review as a conference paper at ICLR 2027

More Multimodality Is Better? Uncovering the Functional Roles of Multimodal Information in Agent Memory

Abstract

Multimodal large language model (MLLM) agents increasingly rely on multimodal memory to store and reuse past experiences for long-horizon tasks, yet it remains unclear when, where, and how multimodal information benefits memory. We present the first systematic study of how different levels of multimodal integration within the memory pipeline affect agent performance. We establish a four-level multimodal memory hierarchy that progressively introduces multimodal capabilities into memory construction, storage and retrieval, and utilization, ranging from text-only to fully multimodal memory. Experiments are conducted under long-horizon settings across 4 representative memory-intensive task domains, 10 benchmarks, and 4 MLLM backbones. Our results reveal that deeper multimodal integration does not consistently improve performance. Instead, its benefits depend on whether newly introduced multimodal capabilities address remaining evidence gaps. Three critical roles of multimodal information are identified: a Semantic Enricher during memory construction, an Experience Locator during retrieval, and an Evidence Resolver during utilization—capturing distinct ways multimodal information supports downstream decision-making. Deeper integration, however, may also introduce distraction and yield limited or negative returns. These findings challenge the prevailing assumption that richer multimodal memories universally benefit agent performance and highlight the importance of stage-aware multimodal memory design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.