acceptodds
Under review as a conference paper at ICLR 2027

RSIBench-Code: Can LLM Agents Handle Enterprise-Level LLM Development Code for Self-Improvement?

Abstract

Recursive self-improvement (RSI) aims to enable AI systems to contribute to the processes that produce stronger future models. A practical route is model self-training, but the code behind training, evaluation, serving, and infrastructure is still largely designed, maintained, and debugged by human engineers. This raises a central question: **can LLM agents handle the enterprise-level LLM-development code needed for self-improving AI**? We introduce **RSIBench-Code**, a benchmark for evaluating LLM agents on coding tasks reconstructed from practical LLM development problems in an enterprise. The benchmark contains 105 tasks across post-training, training infrastructure, serving, evaluation, and agent systems. These tasks cover implementation, repair, test writing, controlled experiments, and execution-path alignment. In each task, agents must locate and modify relevant code within repositories containing approximately 1.8 million source lines, with changing up to thousands of lines. RSIBench-Code supports long-horizon coding attempts with up to 150 agent turns and a 7,200-second solving budget. Experiments show that existing agents make meaningful but incomplete progress, suggesting that enterprise-level LLM-development code remains a key bottleneck for recursive self-improvement. Furthermore, RSIBench-Code provides a continuing source of executable evaluation tasks and potential training data for RSI in future work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.