acceptodds
Under review as a conference paper at ICLR 2027

Vantage: Learning Memory Advantage by Asymmetric Re-Attempt

Abstract

Agentic memory enables LLM agents to self-evolve without weight updates by accumulating experience in a memory bank and retrieving it for new queries. As the bank grows, the key is to learn each memory's utility score so that retrieval prioritizes the most useful memories. Existing methods learn utility from the outcome (success or failure) of each attempt, crediting a memory whenever it is retrieved in a successful attempt, even when the attempt would succeed without it. We therefore argue that utility should also account for memory advantage, i.e., the outcome change that a memory brings compared to a memory-free attempt on the same query. To this end, we propose Vantage, an efficient utility-learning framework that incorporates outcome changes without the cost of memory-free attempts: the change is approximated by immediately re-attempting the query with the updated memory bank, so the first attempt itself serves as the reference. The re-attempt is asymmetric: only failed queries are re-attempted, whereas successful queries, whose outcome rarely changes, are not; this limits repeated crediting and redundant memories. Under frozen-bank evaluation on diverse tasks spanning coding, computer use, and embodied planning, Vantage improves success rate (SR) over the strong baseline MemRL by an absolute 6.4% on training queries and 2.2% on held-out queries, with 40% fewer attempts and thus a 40% lighter memory bank. Incorporating outcome changes into MemRL also improves it, further validating the advantage view. Detailed ablations confirm that both the asymmetric re-attempt and the incorporation of outcome changes drive the gains, and the learned utilities align more closely with memory advantage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.