RSIBench: Can AI Agents Improve Their Own Harnesses?
Abstract
Recursive self-improvement (RSI) promises a simple yet powerful loop: an AI system improves itself, and the modified version continues the process. As agent harnesses increasingly shape how language models use context, tools, and feedback, evolving those harnesses offers a plausible near-term way to realize this loop. But can today's agents improve their own harnesses? Do these gains generalize to unseen tasks? Which harness changes help? To answer these questions, we introduce RSIBench, a benchmark that tests whether agents can improve their own harnesses. At each iteration, an agent uses task trajectories to revise its open-source Pi harness , while the model weights remain fixed. If a harness revision improves performance on training tasks, it replaces the current harness. The agent then uses the revised harness to solve tasks and propose the next revision. Unlike most prior work, we evaluate the final harness on held-out tasks. We evaluate seven models on 100 tasks across Terminal, Environment Learning, SWE, and Science. We define RSI lift as the final harness's held-out score minus the initial harness's held-out score. Claude Opus 5 achieves an RSI lift of +8.3 percentage points (36.4→44.7) across all held-out tasks. However, for GPT-5.5 the gains are less clear: it improves on Science training tasks but loses 8.3 points on held-out tasks (11.1→2.8), and cross-domain transfer is uneven, underscoring the importance of evaluating RSI on held-out tasks. Overall, RSIBench tests a concrete step toward RSI: whether a model can improve its own harness and, crucially, generalize those gains to unseen tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.