Learning to Access Multimodal Agent Memory: Agentic Retrieval with State-Level On-Policy Self-Distillation
Abstract
Long-term multimodal memory enables agents to leverage temporally distributed evidence across interleaved conversations and images. Nevertheless, prevailing approaches largely rely on fixed retrieval strategies or rule-based agentic workflows, offering limited flexibility to adapt retrieval decisions as the available evidence evolves. Effective evidence acquisition requires evaluating evidence sufficiency and identifying unresolved information gaps before taking the next action. Accordingly, we propose state-level on-policy self-distillation (OPSD), in which privileged information (PI) explicitly captures these two factors. At each intermediate state of the student rollout, a verifier reconstructs these diagnoses by comparing the current short-term memory (STM) with the reference answer, providing privileged context for a frozen teacher to guide subsequent memory-access decisions. OPSD then aligns the student’s action preferences with those of the privileged teacher, enabling evidence-dependent decisions on when to terminate and what evidence to acquire next without privileged information at inference time. To support this learning formulation, we develop a unified agentic memory-access framework that formulates evidence acquisition as sequential tool use, provides heterogeneous memory-access operations, and maintains a compact STM for consolidating intermediate evidence. We further construct MemVista, a training dataset comprising long-horizon multimodal conversations with temporally evolving facts and traceable supporting evidence. Experiments on Mem-Gallery demonstrate that our approach improves the mean LLM-judge score from 0.780 to 0.852 compared with the strongest external baseline. These results indicate that state-level OPSD with dynamic PI enables more effective memory access and substantially improves performance on long-term multimodal memory reasoning tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.