acceptodds
Under review as a conference paper at ICLR 2027

Recursive Experiential–Working Memory Evolution for Long-Horizon Agent Harnesses

Abstract

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation, while failures are difficult to trace to specific harness components, hindering effective evolution. We introduce Recuris, a recursive Experiential–Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes failures to specific memory components. Across tasks, a fixed Meta-Agent uses this evidence to make localized, validation-gated updates to the memory system. These updates reshape execution and yield new evidence, forming a bounded recursive memory-evolution loop. Across four challenging long-horizon benchmarks and ten models, substantially improves task success for both frontier and open-source models: on -Bench Retail it adds points to GPT-5.6 Sol and to Claude Opus 5, taking Opus 5 to , and / points on Qwen3.6-27B/35B on SkillFlow. The advantage widens as the interaction horizon grows, to points on the longest tasks, and common long-horizon failures fall by up to . These results position recursively evolving memory as a scalable foundation for RSI, enabling agents to continuously transform accumulated experience into increasingly effective long-horizon behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.