Credit-Aware Memory Policy Optimization for Self-Evolving Black-Box Agents
Abstract
Self-evolving agents improve online by reusing experience accumulated in a memory of past user interactions. Adapting the agent itself from deployment feedback requires model parameters that the strongest API-served models do not expose, so the practical alternative learns per-memory utilities that decide which memories are retrieved for the current state. Such utilities, however, are updated with a shared trajectory-level return, which produces Memory Credit Leakage: because co-retrieved memories are indistinguishable to a single outcome signal, irrelevant and even harmful entries inherit the credit of the useful memories retrieved with them and accumulate spurious reinforcement as deployment lengthens. We argue that a memory deserves retrieval only if including it improves the return over excluding it, a contrast that no shared return can reveal. Therefore, we propose Credit-Aware Memory Policy Optimization (CMPO), an online framework that recovers this contrast from a single rollout by randomizing what the black-box agent sees rather than how it acts. CMPO exposes each memory through a stochastic gate with a recorded inclusion probability, so one rollout already contains both included and excluded memories, and reweighting the joint return by these probabilities yields a per-memory comparison that updates the gate, exposing useful memories more often and suppressing harmful ones. We prove that this credit is an unbiased estimator of the gate-local inclusion–exclusion return contrast and equals the derivative of the gate-local objective with respect to the inclusion probability. Across three benchmarks covering terminal interaction, embodied planning, and tool use, CMPO ranks first in 17 of 24 white-box and 20 of 24 black-box settings, sustaining accurate credit attribution over long-horizon deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.