acceptodds
Under review as a conference paper at ICLR 2027

RSI-Index: Measuring Autonomous LLM R&D in Frontier Agent Systems

Abstract

Recursive self-improvement requires AI systems that can do AI research autonomously, but this capability is measured primarily by model developers, on their own models and internal tasks. We introduce RSI-Index, an external benchmark that measures how much of the research behind developing LLMs an agent can do on its own. Its four tasks cover pre-training, post-training and harness engineering, and each scores the agent's result against the best known result. Across 19 models from eight providers, the strongest, Claude Opus 5.5, scores 0.373. Performance has risen steadily across model releases, although gains so far come mainly from combining established techniques. If current trends continue and generalize to frontier development, within a year, autonomous AI agents could drive improvements in frontier language-model performance comparable to major human-led advancements.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.