AdaM-TTT: Adaptive Agent Memory via Test-Time Training
Abstract
Long-horizon language agents must continually adapt as tasks evolve during interaction. Agent memory supports adaptation at deployment by making relevant past experience available during interaction. Existing methods, however, operate mainly through contextual memory—retrieved trajectories or summaries placed in the current prompt—while the policy parameters remain fixed. This confines adaptation to information repeatedly supplied in the prompt, even when useful guidance should persist across subsequent decisions. We investigate whether retrieved experience can also supervise online policy adaptation, allowing memory to influence both the agent's context and its parameters. Although model parameters are conventionally viewed as long-term memory, we use a lightweight adapter as short-term parametric memory: it is updated within an episode, replaced when guidance changes, and never modifies the frozen base model. We introduce AdaM-TTT (Adaptive Agent Memory via Test-Time Training), a framework that retrieves successful experiences and distills them into a compact procedural memo relevant to the agent's current task state. The procedural memo guides action generation as contextual memory; together with a compact live-progress summary, it forms the next-token-prediction target used to update a LoRA adapter at test time. This combines immediate contextual guidance with temporary knowledge encoded in the adapter. The policy can retain this paired memory state, update it with new experience, or reset it by clearing the memo and unloading the adapter, allowing adaptation to remain responsive as the episode progresses while the base model stays frozen. On ALFWorld, AdaM-TTT achieves a 5.0% relative improvement over the strongest memory baseline, while test-time training contributes only 2.5% of estimated parameter-level compute. These results demonstrate the benefit of coupling contextual memory with short-term parametric adaptation, providing a new mechanism for continual adaptation from experience after deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.