RSIE: Recursive Self-Improvement through Interaction with Expert Environments
Abstract
Recent work on recursive self-improvement (RSI) has demonstrated that agents can improve through repeated interaction with their tasks and environments. Yet it remains unclear whether such improvement can persist in expert environments, where executions are long-horizon and costly, and feedback is sparse, delayed, and domain-specific. Research and professional tasks naturally exhibit these characteristics, making them a demanding testbed for recursive self-improvement. We introduce RSIE, a system for recursive self-improvement through interaction with expert environments. RSIE repeatedly executes expert tasks, extracts reusable experience from completed trajectories and their evaluation feedback, and retains candidate experience only when its reuse leads to measurable improvement. Confirmed experience is accumulated through two complementary paths: externalized guidance for subsequent executions and parameter internalization via on-policy self-distillation. As the agent improves, saturated tasks are replaced by new ones, continually exposing the system to its evolving capability frontier. We evaluate RSIE with Qwen3.6-35B and Qwen3.5-9B on ResearchClawBench, SGIBench, and GDPVal, spanning scientific research and professional work. RSIE consistently improves over the corresponding base agents, showing that recursive self-improvement can extend to expert environments with sparse task-level feedback.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.