Experience-Handoff Reinforcement Learning for Recursive Self-Improvement
Abstract
It is common to train Large Language Models (LLMs) for agentic tasks with Reinforcement Learning (RL) from scalar rewards. However, this throws away the fine-grained information about the environment that the agent experienced during the trial (e.g., error logs). To utilize this rich information, recent works wrap the agent in a harness that refines its own inputs through reflection or compaction and hands past experience to future rollouts, but the construction of the handed-off artifacts (e.g., a playbook) is itself not learned. We want the best of both worlds: gradient updates from a scalar reward, while also utilizing environment information to hand off experience. As a concrete testbed, we fix the Agentic Context Engineering () harness and train one model to act, reflect, and curate task-local playbooks. Our method, , evaluates candidate experience-handoff artifacts by their downstream effect on future actor rollouts, and co-optimizes the role-specific behaviors (action generator, insight reflector, and playbook curator) under a fixed scaffold. Across different agentic benchmarks, shows faster convergence and higher converged performance than existing baselines, including outcome-only . The results suggest that agent training should optimize not only the actions in the current rollout, but also the model-written experience that future rollouts inherit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.