EvoRelay: Asynchronous Online Experience Learning for Long-Horizon Agents
Abstract
Large language model (LLM) agents increasingly solve long-horizon tasks in interactive environments, where past interactions provide reusable experience for future decisions. Existing memory-based self-evolution methods commonly organize experience at the task level, retrieving historical experience by global task similarity and making newly acquired experience available only after task completion. However, long-horizon tasks typically unfold through functional stages with distinct objectives. This task-level design creates both a granularity mismatch and an efficiency bottleneck: global similarity may fail to identify experience relevant to the current stage, while newly acquired experience cannot benefit subsequent tasks until the source task completes, limiting the efficiency of online experience learning. To address these limitations, we propose Evorelay, an asynchronous online experience learning framework that unifies experience retrieval, extraction, and cross-task propagation around functional stages. Within this unified organization, historical experience is retrieved according to the current stage objective, and experience learned from completed stages becomes available to subsequent tasks while the source task continues. This asynchronous relay reduces cross-task experience propagation latency and substantially improves the efficiency of online experience learning. Across four backbone models, Evorelay improves task success over Vanilla Agent by an average of 6.0 percentage points on SWE-bench Verified and 3.8 percentage points on AppWorld. It achieves the highest average task success across the evaluated settings, with 2.6–3.4 mean end-to-end speedups over two representative memory-based self-evolution baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.