Joint Agent Memory and Exploration Learning via Novelty Signals
Abstract
Exploration and memory are two fundamental capabilities for autonomous agents in open-ended environments. However, both are difficult to learn. Effective exploration requires reasoning about the possible consequences of actions and identifying actions that can lead to new states, while a good memory system needs to compress large-scale interaction histories into a form that is easy for an agent to use. Instead of considering these two capabilities separately, we observe that memory and exploration form a mutually dependent loop: sustained exploration requires memory to distinguish exhausted behaviors from unseen ones, while novelty-seeking interactions provide supervision for learning useful memory representations. Based on this insight, we introduce **J**oint **A**gent **M**emory and **E**xploration **L**earning (**JAMEL**), a framework that jointly trains agentic memory and an exploration policy through novelty-driven interaction. JAMEL compresses interaction histories into latent memory tokens and jointly trains memory and policy through novelty-guided rejection sampling fine-tuning on filtered exploration prefixes. We establish evaluation settings for open-environment exploration across GUI, navigation, and household domains. Across these domains, JAMEL consistently outperforms the evaluated baselines, supporting the effectiveness of jointly learning memory and exploration through novelty signals.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.