acceptodds
Under review as a conference paper at ICLR 2027

RSI-Env: An Open-Ended Environment for Recursive Self-Improvement of Agents

Abstract

Autonomous agents can execute complex tasks, but their performance depends on the underlying model, context-management design, available tools, and execution framework. The vision of recursive self-improvement (RSI) is for agents to autonomously improve their own models and harnesses within an open-ended search space, allowing improved successor systems to continue the process. Existing methods typically optimize predefined components under fixed objectives and evaluation procedures, leaving limited support for system-level exploration and cross-generation inheritance. We present RSI-Env, a persistent sandbox for open-ended system exploration, with an independent host controller managing evaluation, services, and versioning. Three settings, Benchmark-Guided, Capability-Directed, and Open-General, vary goal specification and feedback to assess research autonomy. Starting from Qwen3.8-27B, we evaluate Self and a fixed GPT-5.6 Sol Teacher across seven task settings. Among the 14 resulting model–harness pairs, 64.3% improve on at least one benchmark, but all regress on at least one benchmark; only one achieves a positive mean change across the four benchmarks. Agents can modify systems, construct research materials, and perform engineering and capability validation, but capability evidence only occasionally guides decisions. We identify three structural failures: insufficient capability validation, misalignment between evidence and update selection, and failure to retain local gains. Sustained and independently verified capability improvement remains unresolved.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.